Aller au contenu principal
Accès ouvert déclaré 2025 preprint

GhostBuster: A Deep-Learning-based, Literature-Unbiased Gene Prioritization Tool for Gene Annotation Prediction

0Citations signalées, ce qui n’est pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : gb. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Abstract All genes are not equal before literature. Despite the explosion of genomic data, a significant proportion of human protein-coding genes remain poorly characterized (“ghost genes”). Due to sociological dynamics in research, scientific literature disproportionately focuses on already well-annotated genes, reinforcing existing biases (bandwagon effect). This literature bias often permeates machine learning (ML) models trained on gene annotation tasks, leading to predictions that favor well-studied genes. Consequently, standard ML performance metrics may overestimate biological relevance by overfitting literature-derived patterns. To address this challenge, we developed GhostBuster, an encoder-decoder ML platform designed to predict gene functions, disease associations and interactions while minimizing literature bias. We first compared the impact of biased (Gene Ontology) versus unbiased training datasets (LINCS, TCGA, STRING). While literature-biased sources yielded higher ML metrics, they also amplified bias by prioritizing well-characterized genes. In contrast, models trained on unbiased datasets were 2-3× more effective at identifying recently discovered gene annotations. Notably, one of the unbiased channels (TCGA), combined minimal amounts of literature bias with robust performance, at a test ROC-AUC of 0.8-0.95. We demonstrate that GhostBuster can be applied to predict novel gene functions, refine pathway memberships, and prioritize intergenic GWAS hits. As the first ML framework explicitly designed to counteract literature bias, GhostBuster offers a powerful tool for uncovering the roles of understudied genes in cellular function, disease, and molecular networks. Graphical Abstract GCN: graph convolutional neural network; GeneOntol.: Gene Ontology; LINCS: Library of Integrated Network-Based Cellular Signatures; MLP: multi-layer perceptron (deep learning); PPI: (physical) protein-protein interactions; SVM: support vector machine; TCGA: The Cancer Genome Atlas; y/n: yes or no (binary classification).

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
GhostBuster: A Deep-Learning-based, Literature-Unbiased Gene Prioritization Tool for Gene Annotation Prediction
Date Crossref
27/06/2025
Éditeur
openRxiv
Type
posted-content

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Bioinformatics and Genomic NetworksMachine Learning in BioinformaticsGene expression and cancer classification

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.