Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database
Rattachement africain : jp, Afrique du Sud. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
BioSample is a repository of experimental sample metadata. It is a comprehensive archive that enables searches of experiments, regardless of type. However, there is substantial variability in the submitted metadata due to the difficulty in defining comprehensive rules for describing them and the limited user awareness of best practices in creating them. This inconsistency poses considerable challenges to the findability and reusability of archived data. Given the scale of BioSample, which hosts over 40 million records, manual curation is impractical. Automatic rule-based ontology mapping methods have been proposed to address this issue, but their effectiveness is limited by the heterogeneity of the metadata. Recently, large language models (LLMs) have gained attention in natural language processing and are promising tools for automating metadata curation. In this study, we evaluated the performance of LLMs in extracting cell line names from BioSample descriptions using a gold-standard dataset derived from ChIP-Atlas, a secondary database of epigenomics experiment data in which samples were manually curated. The LLM-assisted methods outperformed traditional approaches, achieving higher accuracy and coverage. We further extended them to extract information about experimentally manipulated genes from metadata when manual curation had not yet been applied in ChIP-Atlas. This also yielded successful results, including the facilitation of more precise filtering of the data and the prevention of possible misinterpretations caused by the inclusion of unintended data. These findings underscore the potential of LLMs in improving the findability and reusability of experimental data in general, which would considerably reduce the user workload and enable more effective scientific data management.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database
- Date Crossref
- 01/01/2025
- Éditeur
- Oxford University Press (OUP)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Hiroshima University Genome Editing Innovation Center pays non établi dans la noticeUniversité ou école supérieure
-
Research Organization of Information and Systems pays non établi dans la noticeStructure de recherche
-
Database Center for Life Science pays non établi dans la noticeStructure de recherche
-
The University of Tokyo pays non établi dans la noticeUniversité ou école supérieure
-
Kumamoto University Institute of Resource Development and Analysis pays non établi dans la noticeUniversité ou école supérieure
-
Research Research (South Africa) Afrique du Sud (code pays fourni par la source)Entreprise
-
Chiba University Institute for Advanced Academic Research pays non établi dans la noticeUniversité ou école supérieure
-
Graduate School of Integrated Sciences for Life pays non établi dans la noticeUniversité ou école supérieure
-
BioData Science Initiative pays non établi dans la noticeInstitution
-
Graduate School of Medicine pays non établi dans la noticeUniversité ou école supérieure
Genome Editing Innovation Center — Hiroshima University, Research Organization of Information and Systems et Database Center for Life Science, avec 7 autres affiliations. Pays d’affiliation : Afrique du Sud.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.