Aller au contenu principal
Accès ouvert déclaré 2026 preprint

Accurate ab initio gene prediction in eukaryotes with Tiberius in multiple clades

0Citations signalées — pas une note de qualité
9Institutions déclarées
3Pays d’affiliation déclarés

Résumé fourni par la source

Abstract Background Eukaryotic genome annotation is currently bottlenecked by limitations in the generality, scalability or accuracy of computational methods and fewer than 20% of the genomes available at NCBI Datasets have associated gene annotations. Evidence-based pipelines such as BRAKER3 achieve high accuracy but require substantial extrinsic evidence and compute. Deep learning approaches have recently achieved large improvements in ab initio gene prediction accuracy. The end-to-end deep learning gene predictor Tiberius approaches the accuracy of evidence-based annotation on mammalian genomes without using any extrinsic evidence, but its published models were trained on mammals only, limiting its applicability across eukaryotes. Results We extend Tiberius beyond mammals by training lineage-specific models for Mesan-giospermae, Fungi, Vertebrata, Insecta, Chlorophyta and Bacillariophyta, making the tool applicable to 92% of currently available eukaryotic assemblies. Across a benchmark of 33 species, Tiberius achieved higher exon-, transcript- and gene-level accuracy than the other evaluated ab initio methods, Helixer and ANNEVO, improving gene-level F1 score by 12–37 percentage points over Helixer and by 10–22 percentage points over ANNEVO, while also having the fastest runtimes overall. Compared with BRAKER3, which incorporates RNA-Seq and protein evidence, Tiberius approaches state-of-the-art accuracy in Mesangiospermae, Fungi, Bacillariophyta and Chlorophyta, while being on average 80 times faster when using a GPU. A reimplementation of the Tiberius backend reduced runtime by 31% compared to the previous implementation. Tiberius and its models are also available through a web server, which allows annotation of submitted assemblies without local resources. Furthermore, the Vertebrata model has been applied to annotate 2,948 vertebrate assemblies totalling nearly 6 trillion base pairs. Conclusions Tiberius is transferable well beyond Mammalia and reaches accuracy close to evidence-based annotation in several clades at a small fraction of the compute cost. This makes it a practical choice for highly accurate large-scale genome annotation, particularly when extrinsic evidence is unavailable or annotation throughput is limiting. Availability and implementation Code: https://github.com/Gaius-Augustus/Tiberius , web server: https://bioinf.uni-greifswald.de/tiberius .

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé, mais le titre doit être comparé manuellement.

Titre Crossref
Accurate <i>ab initio</i> gene prediction in eukaryotes with Tiberius in multiple clades
Date Crossref
28/04/2026
Éditeur
openRxiv
Type
posted-content

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Genomics and Phylogenetic StudiesProtist diversity and phylogenyMachine Learning in Bioinformatics

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, ROR et la Banque mondiale, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune donnée externe enregistrée en base. Sources et limites.