Aller au contenu principal
Accès ouvert déclaré 2026 preprint

The StrainDiscoveryDatabase: an open framework for standardized microbial strain data

0Citations signalées, ce qui n’est pas une note de qualité
6Institutions déclarées
4Pays d’affiliation déclarés

Rattachement africain : de, us, es, nl. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Abstract The vast amount of existing data on microbial strains holds immense potential to revolutionize bioindustry through the application of Artificial Intelligence (AI). However, the training of robust predictive AI models requires large-scale, unified, and non-redundant microbial datasets, which is currently severely hindered by the deep fragmentation of the data and the existence of synonymous strain identifiers in different culture collections. To overcome these infrastructural bottlenecks, we have established the StrainDiscoveryDatabase (SDD), a comprehensive, machine-readable dataset encompassing over 6.2 million harmonized data points for 256,889 microbial strains. The SDD does not rely on its own data repository, but rather on existing data that is retrieved on the fly from highly curated databases. Through an automated pipeline phenotypic, genotypic, and contextual data are systematically retrieved via the Application Programming Interfaces (APIs) of the Bacterial Diversity database (Bac Dive ), the Microbial Resource Research Infrastructure Information System (MIRRI-IS) and the catalogue of the DSMZ. In order to reliably resolve synonymous strain identifiers, the StrainInfo database and its identification tools are employed, enabling the accurate deduplication and unification of records from disparate sources. The resulting aggregated, globally unique dataset is provided in a highly standardized JSON format in strict adherence to the FAIR data principles. By bridging isolated database silos and linking distributed knowledge to discrete biological entities, the SDD provides a high-quality, foundational resource designed to accelerate trait-based strain discovery, large-scale comparative analysis, and machine learning applications, thereby supporting the translation of the extensive existing knowledge on microbial traits to bioindustrial applications.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
The StrainDiscoveryDatabase: an open framework for standardized microbial strain data
Date Crossref
10/09/2026
Éditeur
openRxiv
Type
posted-content

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Bacterial Identification and Susceptibility TestingGenomics and Phylogenetic StudiesGut microbiota and health

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.