GUANinE v1.1 reveals complementarity of supervised and genomic language models
Rattachement africain : us. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
There has been much debate about the benefits of supervised versus unsupervised learning on genomes. Determining which is better in what contexts requires developing comprehensive benchmarks spanning functional and evolutionary tasks. Importantly, such benchmarks need large sample sizes to enable well-powered ranking of models. Having developed and applied such a benchmark here (GUANinE v1.1), we conclusively demonstrate each paradigm offers key advantages and outperforms on certain tasks. In accordance with training, supervised sequence-to-function models exhibit strong performance when annotating functional states characterized by chromatin accessibility or histone marks, while self-supervised language models outperform on evolutionary conservation. Our hundreds of new evaluations in this v1.1 expansion provide evidence for a tradeoff between input context size and model parameter count for a fixed compute budget, which we depict with new metrics such as kiloparameters/base pair. We also construct two new large-scale variant interpretation tasks in v1.1: cadd-snv measuring deleteriousness, and clinvar-snv measuring clinical pathogenicity. We find that conservation scores, and by extension, genomic language models, predict deleteriousness well, but successfully translating deleteriousness predictions to pathogenicity remains challenging. GUANinE v1.1 newly evaluates dozens of pretrained genomic models, and we conclude that moderate-context hybrid or post-trained language models may define the next era of machine learning in genomics.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- GUANinE v1.1 reveals complementarity of supervised and genomic language models
- Date Crossref
- 01/08/2026
- Éditeur
- Oxford University Press (OUP)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Berkeley College pays non établi dans la noticeUniversité ou école supérieure
-
University of California pays non établi dans la noticeUniversité ou école supérieure
-
Chan Zuckerberg Initiative (United States) pays non établi dans la noticeEntreprise
-
Santa Cruz County Office of Education pays non établi dans la noticeOrganisme public
-
Center for Computational Biology pays non établi dans la noticeInstitution
-
UC Berkeley pays non établi dans la noticeInstitution
-
Chan Zuckerberg Biohub pays non établi dans la noticeInstitution
-
Department of Applied Mathematics pays non établi dans la noticeInstitution
-
Department of Electrical Engineering and Computer Sciences pays non établi dans la noticeInstitution
Berkeley College, University of California et Chan Zuckerberg Initiative (United States), avec 6 autres affiliations.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.