Aller au contenu principal
Accès ouvert déclaré 2025 article

Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database

14Citations signalées, ce qui n’est pas une note de qualité
7Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : jp, Afrique du Sud. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

BioSample is a repository of experimental sample metadata. It is a comprehensive archive that enables searches of experiments, regardless of type. However, there is substantial variability in the submitted metadata due to the difficulty in defining comprehensive rules for describing them and the limited user awareness of best practices in creating them. This inconsistency poses considerable challenges to the findability and reusability of archived data. Given the scale of BioSample, which hosts over 40 million records, manual curation is impractical. Automatic rule-based ontology mapping methods have been proposed to address this issue, but their effectiveness is limited by the heterogeneity of the metadata. Recently, large language models (LLMs) have gained attention in natural language processing and are promising tools for automating metadata curation. In this study, we evaluated the performance of LLMs in extracting cell line names from BioSample descriptions using a gold-standard dataset derived from ChIP-Atlas, a secondary database of epigenomics experiment data in which samples were manually curated. The LLM-assisted methods outperformed traditional approaches, achieving higher accuracy and coverage. We further extended them to extract information about experimentally manipulated genes from metadata when manual curation had not yet been applied in ChIP-Atlas. This also yielded successful results, including the facilitation of more precise filtering of the data and the prevention of possible misinterpretations caused by the inclusion of unintended data. These findings underscore the potential of LLMs in improving the findability and reusability of experimental data in general, which would considerably reduce the user workload and enable more effective scientific data management.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database
Date Crossref
01/01/2025
Éditeur
Oxford University Press (OUP)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Hiroshima University Genome Editing Innovation Center pays non établi dans la notice
    Université ou école supérieure
  • Research Organization of Information and Systems pays non établi dans la notice
    Structure de recherche
  • Database Center for Life Science pays non établi dans la notice
    Structure de recherche
  • The University of Tokyo pays non établi dans la notice
    Université ou école supérieure
  • Kumamoto University Institute of Resource Development and Analysis pays non établi dans la notice
    Université ou école supérieure
  • Research Research (South Africa) Afrique du Sud (code pays fourni par la source)
    Entreprise
  • Chiba University Institute for Advanced Academic Research pays non établi dans la notice
    Université ou école supérieure
  • Graduate School of Integrated Sciences for Life pays non établi dans la notice
    Université ou école supérieure
  • BioData Science Initiative pays non établi dans la notice
    Institution
  • Graduate School of Medicine pays non établi dans la notice
    Université ou école supérieure

Genome Editing Innovation Center — Hiroshima University, Research Organization of Information and Systems et Database Center for Life Science, avec 7 autres affiliations. Pays d’affiliation : Afrique du Sud.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Biomedical Text Mining and OntologiesSemantic Web and OntologiesGenomics and Phylogenetic Studies

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.