Aller au contenu principal
Accès ouvert déclaré 2026 dataset

Horizyn-1 model weights and development dataset: Dual-encoder contrastive learning accelerates enzyme discovery

0Citations signalées, ce qui n’est pas une note de qualité
0Institutions déclarées
0Pays d’affiliation déclarés

Le résumé fourni par la source

Overview This repository provides the official model weights for the Horizyn-1 enzyme discovery models, as well as the open-source development dataset, as described in "Dual-encoder contrastive learning accelerates enzyme discovery." The accompanying code for training, inference, and evaluation is available at: https://github.com/dayhofflabs/horizyn. Model Checkpoints Two model checkpoints are provided in this archive: horizyn_v1_0_inf.ckpt (Inference Model): The primary model recommended for inference. Note: This model was trained on a much larger, withheld dataset and was not trained on the development dataset provided in this repository. horizyn_v1_0_dev.ckpt (Development Model): A model trained entirely on the accompanying open-source development dataset, provided for reproducibility and benchmarking. Development Dataset Overview & Splits The dataset included in this repository was used to train and evaluate the development model (horizyn_v1_0_dev.ckpt). It includes the full set of reaction SMILES, protein-reaction pairs, and pre-computed ProtT5 protein embeddings. The training and test sets were created by splitting on reactions to prevent data leakage. The test set was strictly filtered to exclude any reactions with high similarity to any reactions in the training set (see manuscript for details). For evaluation, the test setup involves identifying the correct enzyme for each test reaction from a total screening pool of 216,132 proteins contained in this dataset. Dataset Statistics Reactions: 10,785 (Train) / 1,012 (Test) Enzymes: 192,769 (Train) / 32,100 (Test) Reaction-Enzyme Pairs: 257,733 (Train) / 33,996 (Test) File Manifest The archive contains the following standardized files: horizyn_v1_0_inf.ckpt & horizyn_v1_0_dev.ckpt Model weights for the Horizyn-1 inference and development models, respectively. train_rxns.csv & test_rxns.csv Contains reaction SMILES strings for the development dataset. Columns: rs_id, reaction_id, reaction_smiles train_pairs.csv & test_pairs.csv Defines the positive training and testing pairs for the development dataset. Columns: pr_id, reaction_id, protein_id prots_t5.h5.gz Compressed HDF5 file containing pre-computed protein embeddings (ProtT5-XL). Structure (when uncompressed): /ids: Dataset of protein IDs (strings) /vectors: Dataset of embeddings (float32, shape: [N, 1024]) prots.fasta.gz (Reference) Compressed FASTA file containing protein sequences corresponding to the embeddings in the HDF5 file. Citation If you use these models or the development dataset, please cite the associated publication: Rocks, J. W., Truong, D. P., Rappoport, D., Maddrell-Mander, S., Martin-Alarcon, D. A., Lee, T. M., Crossan, S., & Goldford, J. E. (2026). Dual-encoder contrastive learning accelerates enzyme discovery. Proc. Natl. Acad. Sci. U.S.A. 123 (12) e2520070123. https://doi.org/10.1073/pnas.2520070123

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

La source scientifique ouverte est momentanément indisponible.

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.