19. Expanding MAVE data maps for use in human genomics applications
Rattachement africain : us, gb, au. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Introduction With the growing use of high-throughput sequencing technologies in the clinical setting, there is a growing demand for functional evidence supporting variant interpretation in clinically impactful genes. In recent years, multiplexed assays of variant effect (MAVEs) have been introduced as a new line of functional evidence to support variant classification in a growing number of genes, allowing for the functional effects of up to thousands of variants to be investigated in parallel. In 2023, we first mapped MAVE variants in MaveDB to human reference sequences, generating a set of variant mappings for 2.5 million protein and genomic variants across 207 datasets. Since producing this initial mapping set, new datasets have continued to be added to MaveDB, necessitating updates to our workflow to support the increased scale, complexity, and diversity of these new data. Here, we describe these improvements and how they have enabled the dissemination of over 1,057 MAVE score sets for use in human genomics research. Methods Our variant mapping workflow, consisting of MAVE sequence alignment, reference sequence selection, and variant processing using the GA4GH Variation Representation Specification, was developed and released as the open-source dcd-mapping Python package. Using this software, we translated approximately 9.0 million protein and genomic variants across 1,057 score sets from MaveDB. The updated mapped dataset was then uploaded to a publicly accessible s3 bucket, enabling dissemination to downstream implementers. We assessed these data for concordance between experimental and aligned reference sequences. Mapped variants were integrated into common tools and resources for human genomics research. Results We observed that 68.44% (6,158,451/8,998,024) of mapped MAVE variants were concordant, with the majority of discordant variants arising due to mutagenized MAVE sequences that differed from the human reference sequences. The updated variant mapping set was successfully integrated into the Genomics 2 Proteins Portal, the UCSC Genome Browser, Ensembl Variant Effect Predictor, and DECIPHER, and we demonstrated unique applications of this data in each of these resources. Discussion and Conclusion Our updated workflow has enabled variant mapping for hundreds of new score sets in MaveDB, ensuring that MAVE data can continue to be provided to human genomics applications for downstream interpretation. This allows for MAVE data to be linked with other knowledge in various genomics portals, providing additional evidence to assist in variant assessment. Our analysis also highlights challenges in using MAVE data to perform clinical variant classification, specifically in assessing variants that occur in mutagenized regions. By providing a diverse dataset of mapped MAVE score sets, we have helped provide a foundation for the development of expert community-driven guidelines for the clinical application of MAVE data.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- 19. Expanding MAVE data maps for use in human genomics applications
- Date Crossref
- 01/12/2025
- Éditeur
- Elsevier BV
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.