Aller au contenu principal
Accès ouvert déclaré 2022 preprint

Welcome to the big leaves: best practices for improving genome annotation in non-model plant genomes

5Citations signalées — pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Résumé fourni par la source

ABSTRACT • Premise of the study Robust standards to evaluate quality and completeness are lacking for eukaryotic structural genome annotation. Genome annotation software is developed with model organisms and does not typically include benchmarking to comprehensively evaluate the quality and accuracy of the final predictions. Plant genomes are particularly challenging with their large genome sizes, abundant transposable elements (TEs), and variable ploidies. This study investigates the impact of genome quality, complexity, sequence read input, and approach on protein-coding gene prediction. • Methods The impact of repeat masking, long-read, and short-read inputs, de novo , and genome-guided protein evidence was examined in the context of the popular BRAKER and MAKER workflows for five plant genomes. Annotations were benchmarked for structural traits and sequence similarity. • Results Benchmarks that reflect gene structures, reciprocal similarity search alignments, and mono-exonic/multi-exonic gene counts provide a more complete view of annotation accuracy. Transcripts derived from RNA-read alignments alone are not sufficient for genome annotation. Gene prediction workflows that combine evidence-based and ab initio approaches are recommended, and a combination of short and long-reads can improve genome annotation. Adding protein evidence from de novo assemblies , genome-guided transcriptome assemblies, or full-length proteins from OrthoDB generates more putative false positives as implemented in the current workflows. Post-processing with functional and structural filters is highly recommended. • Discussion While annotation of non-model plant genomes remains complex, this study provides recommendations for inputs and methodological approaches. We discuss a set of best practices to generate an optimal plant genome annotation, and present a more robust set of metrics to evaluate the resulting predictions.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Welcome to the big leaves: best practices for improving genome annotation in non-model plant genomes
Date Crossref
07/10/2022
Éditeur
openRxiv
Type
posted-content

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Genomics and Phylogenetic StudiesChromosomal and Genetic VariationsRNA and protein synthesis mechanisms

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.