Lost in translation: the pitfalls of Ensembl gene annotations between human genome assemblies and their impact on diagnostics
Résumé fourni par la source
BACKGROUND: Gene models based on GRCh37 human genome assembly are preferred by many international projects over other updated assemblies (GRCh38 and T2T). Discrepant genes (DGs), those recognized as protein coding in the new but not the old assembly, are ignored by several genomic resources and discarded by variant prioritization tools relying on information based on GRCh37. METHODS: We curated a set of Ensembl genes with discrepant annotations between GRCh37 and GRCh38, additionally matching their RefSeq transcripts. Furthermore, we examined their clinical and phenotypic relevance. RESULTS: = 73). We found many clinically relevant genes in this group of neglected genes, and we anticipate that many more will be found relevant in the future. Important additional annotations such as evolutionary constraint metrics are also not calculated for these genes, further relegating them into oblivion. CONCLUSION: For discrepant genes, the inaccurate label of 'non-protein-coding' has relevant ramifications on clinical genetics. Accurate collation of these genes allows for manual curation in clinically relevant scenarios.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Lost in translation: the pitfalls of Ensembl gene annotations between human genome assemblies and their impact on diagnostics
- Date Crossref
- 19/07/2023
- Éditeur
- Informa UK Limited
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.