CanDrivR-CS : A Cancer-Specific Machine Learning Framework for Distinguishing Recurrent and Rare Variants
Rattachement africain : gb. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Abstract Motivation Missense variants play a crucial role in cancer development, and distinguishing between those that frequently occur in cancer genomes and those that are rare may provide valuable insights into important functional mechanisms and consequences. Specifically, if common variants confer growth advantages, they may have undergone positive selection across different patients due to similar selection pressures. Moreover, studies have demonstrated the significance of rare mutations that arise as resistance mechanisms in response to drug treatment. This highlights the importance of understanding the role of both recurrent and rare variants in cancer. In addition to this, most existing tools for variant prediction focus on distinguishing variants found in normal and disease populations, often without considering the specific disease contexts in which these variants arise. Instead, they typically build predictors that generalise across all diseases. Here, we introduce CanDrivR-CS , a set of cancer-specific gradient boosting models designed to distinguish between rare and recurrent cancer variants. Results We curated missense variant data from the International Cancer Genome Consortium (ICGC). Cancer-type-specific models significantly outperformed a baseline pan-cancer model, achieving a maximum leave-one-group-out cross-validation (LOGO-CV) F1 score of up to 90% for CanDrivRSKCM (Skin Cutaneous Melanoma) and 89% for CanDrivR-SKCA (Skin Adenocarcinoma) , compared to 79.2% for the baseline model. Notably, DNA shape properties consistently ranked among the top features for distinguishing recurrent and rare variants across all cancers. Specifically, recurrent missense variants frequently occurred in DNA bends and rolls, potentially implicating regions prone to DNA replication errors and acting as mutational hotspots. Availability and Implementation All training and test data, and Python code are available in our CanDrivR-CS GitHub repository: https://github.com/amyfrancis97/CanDrivR-CS .
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé, mais le titre doit être comparé manuellement.
- Titre Crossref
- <i>CanDrivR-CS</i> : A Cancer-Specific Machine Learning Framework for Distinguishing Recurrent and Rare Variants
- Date Crossref
- 23/09/2024
- Éditeur
- openRxiv
- Type
- posted-content
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
University of Bristol pays non établi dans la noticeUniversité ou école supérieure
University of Bristol.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.