The StrainDiscoveryDatabase: an open framework for standardized microbial strain data
Rattachement africain : de, us, es, nl. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Abstract The vast amount of existing data on microbial strains holds immense potential to revolutionize bioindustry through the application of Artificial Intelligence (AI). However, the training of robust predictive AI models requires large-scale, unified, and non-redundant microbial datasets, which is currently severely hindered by the deep fragmentation of the data and the existence of synonymous strain identifiers in different culture collections. To overcome these infrastructural bottlenecks, we have established the StrainDiscoveryDatabase (SDD), a comprehensive, machine-readable dataset encompassing over 6.2 million harmonized data points for 256,889 microbial strains. The SDD does not rely on its own data repository, but rather on existing data that is retrieved on the fly from highly curated databases. Through an automated pipeline phenotypic, genotypic, and contextual data are systematically retrieved via the Application Programming Interfaces (APIs) of the Bacterial Diversity database (Bac Dive ), the Microbial Resource Research Infrastructure Information System (MIRRI-IS) and the catalogue of the DSMZ. In order to reliably resolve synonymous strain identifiers, the StrainInfo database and its identification tools are employed, enabling the accurate deduplication and unification of records from disparate sources. The resulting aggregated, globally unique dataset is provided in a highly standardized JSON format in strict adherence to the FAIR data principles. By bridging isolated database silos and linking distributed knowledge to discrete biological entities, the SDD provides a high-quality, foundational resource designed to accelerate trait-based strain discovery, large-scale comparative analysis, and machine learning applications, thereby supporting the translation of the extensive existing knowledge on microbial traits to bioindustrial applications.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- The StrainDiscoveryDatabase: an open framework for standardized microbial strain data
- Date Crossref
- 10/09/2026
- Éditeur
- openRxiv
- Type
- posted-content
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.