Evidence and Metadata Dataset for "Large Language Models for Small-Molecule Discovery: Model Taxonomy, Chemical Grounding, and Experimental Evidence
Résumé fourni par la source
This dataset supports the review “Large Language Models for Small-Molecule Discovery: Model Taxonomy, Chemical Grounding, and Experimental Evidence”. It provides a structured, machine-readable evidence base for the literature synthesis, including: (i) a literature inventory and search/screening seed; (ii) model metadata for representative molecular language models, chemistry-specialized large language models, multimodal chemical foundation models, and integrated molecular-design systems; (iii) a study-level experimental evidence matrix distinguishing automated execution from iterative and adaptive experimental feedback; (iv) datasets and benchmark resources spanning molecular properties, reactions, 3D structures, molecule–text data, instruction resources, spectroscopy, and chemistry-LLM evaluation; (v) taxonomy and extraction codebooks; and (vi) source/provenance registries and validation scripts. The evidence cutoff is 4 September 2026. Unreported values are encoded as NR (“not reported with sufficient specificity”); no synthetic or idealized value is treated as reported evidence. The Stage-2 dataset contains 45 core studies with primary or official source locators, 21 structured model/system records, 8 experimental or closed-loop comparator records, 18 dataset/benchmark resources, and 68 provenance/source-registry entries. The dataset is designed to support reproducibility, evidence auditing, model taxonomy, and study-level assessment of experimental maturity in language-centred small-molecule discovery. A final frozen version will replace this draft before publication of the associated review.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.