Genotype-to-Phenotype Associations with Frequented Region Variants
Rattachement africain : us. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
A pangenome represents the entire sequence content and variation of a population. As collections of complete reference quality genomes become more common, so does the prevalence of pangenomes, necessitating the need for scalable computational methods for their analysis. Previously, we developed FindFRs for identifying Frequented Regions in pangenome graphs, where a Frequented Region is a subgraph that is frequently traversed by multiple sequences. In this work, we propose FindFRs3, which is an updated version of FindFRs capable of identifying Frequented Regions with improved runtime and memory efficiency, enabling the analysis of much larger pangenome graphs. In addition, FindFRs3 identifies Frequented Region Variants (the unique subpaths through each region). We demonstrate the utility of these variants by using them as input features for machine learning models that can predict genotype-to-phenotype associations in a large yeast pangenome. Biological insights gained from these variants show that this novel technique allows for a more nuanced and detailed analysis of larger pangenomes.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Genotype-to-Phenotype Associations with Frequented Region Variants
- Date Crossref
- 05/12/2023
- Éditeur
- IEEE
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.