Aller au contenu principal
Accès ouvert déclaré 2026 dataset

Data for: Gene duplication shaped the origin and evolution of the vertebrate olfactory combinatorial code (https://doi.org/10.64898/2026.08.31.747588)

0Citations signalées, ce qui n’est pas une note de qualité
2Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : gr. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Computed data supporting "Gene duplication shaped the origin and evolution of the vertebrate olfactory combinatorial code" (Zirdeli et al. 2026). olfactory_receptor_downstream_probabilities_v1.zip Predicted binding probability for every receptor x odorant pair, as five independent model runs and their per-pair median. Covers the ancestral matrix (432 reconstructed nodes x 754 odorants), the reference-tree matrix (584 sequences x 754 odorants, of which the 433 human rows were analysed) and the held-out M2OR pairs used to evaluate the model. olfactory_receptor_molecule_embeddings_v1.zip MoLFormer-XL embeddings of the 754 odorants, (754, 768), with the index that joins them to the columns of the matrices above. olfactory_receptor_protein_embeddings_pooled_v1.zip Mean-pooled ESM Cambrian 300M representations, (N, 960), for the 722 human class A GPCRs, the 1,399 assayed M2OR sequences and the 584 reference-tree sequences. This is the representation the binding model consumes. olfactory_receptor_chemical_space_v1.zip Derived data for the odorant versus natural-product analysis, built on COCONUT release 08-2026: the analysis table (730,571 molecules x 57 columns - structures, 22 RDKit descriptors, 16 functional-group flags, estimated boiling point, NPClassifier pathway, odorant-set membership), 50 principal components of the MoLFormer embedding, both UMAP projections, the cached COCONUT parent structures and the resolved source-organism table. The 2.1 GB raw 768-dimensional embedding is not included; it is deterministic given the pinned MoLFormer revision and can be regenerated on a GPU. olfactory_receptor_asr_node_runs_v1.zip The 433 per-node IQ-TREE ancestral-state runs as produced, one directory per node. Each contains the node alignment, the treefile, the run log and report, the reconstructed sequence with its gap mask, and the .state file recording the full posterior distribution over all 20 amino acids at every site. Per-residue protein embeddings olfactory_receptor_protein_embeddings_m2or_v1.zip olfactory_receptor_protein_embeddings_gpcr_v1.zip olfactory_receptor_protein_embeddings_reference_tree_v1.zip Per-residue ESM Cambrian 300M arrays (L, 960), one .npy per sequence, for the same three sets. USAGE Clone the code repository, then let it fetch and verify these archives: git clone https://github.com/cgenomicslab/olfactory-receptors.git cd olfactory-receptors conda env create -f environment.yml && conda activate olfactory-receptors python scripts/download_zenodo_data.py # the ~49 MB core python scripts/download_zenodo_data.py --chemicals # or --asr, --embeddings-full, --all python scripts/check_reproducibility.py Each archive's SHA-256 is recorded in zenodo_manifest.json in the repository and is checked before extraction. Every archive carries a destination path, so extraction reproduces the directory layout the analysis code expects; every path in the repository is relative to its root. Files can also be downloaded and unzipped by hand into the destinations listed in that manifest. SOURCES Receptor-odorant bioassays from M2OR (https://m2or.chemsensim.fr/); protein sequences from UniProt; domain models from Pfam via InterPro; natural-product background from COCONUT release 08-2026 (https://coconut.naturalproducts.net); odour descriptors from Pyrfume (https://pyrfume.org), Leffingwell and GoodScents; organism names resolved against NCBI Taxonomy. Analysis code: https://github.com/cgenomicslab/olfactory-receptors

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

La source scientifique ouverte est momentanément indisponible.

Où se fait cette recherche

  • FORTH Institute of Molecular Biology and Biotechnology pays non établi dans la notice
    Structure de recherche
  • FORTH Institute of Applied and Computational Mathematics pays non établi dans la notice
    Structure de recherche

FORTH Institute of Molecular Biology and Biotechnology et FORTH Institute of Applied and Computational Mathematics.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.