Metadata for MetaOrion pretraining samples
Résumé fourni par la source
To enable large-scale representation learning, we assembled a diverse collection of human metagenome datasets spanning multiple body sites and geographical regions. The final dataset comprises 107,494 samples, with the majority originating from Asia, Europe, and North America, and stool samples accounting for approximately 72%. This dataset provides comprehensive metadata for the cohorts utilized during the MetaOrion pretraining phase, including sample unique identifiers (sample_id), anatomical origins (body_site), precise geographical provenances (country and continent), and original publication identifiers (study_name). To ensure maximal transparency and full community reproducibility, we have curated the public repository footprints—specifically including project accessions (project.accession), raw sequencing run IDs (sequencing.data.accession), and their hosting databases (sequencing.data.db)—alongside essential technical specifications and sequencing attributes, such as total sequencing depth (number_bases), instrument models (sequencing_platform), DNA extraction protocols (DNA_extract_type), and the reference database versions utilized for upstream taxonomic profiling (MetaPhlAn4.database).
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.