Aller au contenu principal
Accès ouvert déclaré 2025 dissertation

Holistic Integration of Deep Learning Models for Mass Spectrometry-Based Peptide Identification

0Citations signalées — pas une note de qualité
0Institutions déclarées
0Pays d’affiliation déclarés

Résumé fourni par la source

Tandem mass spectrometry (MS/MS) coupled with liquid chromatography and/or ion mobility separation is an essential experimental technique for high-throughput discovery of the proteome, or the collection of proteins in a biological system. Computational tools have revolutionized the field, automating the grueling manual annotation of MS/MS spectra with the peptides that produced them. Assigning peptide sequences to spectra is error-prone even for software solutions, requiring statistical methods to limit the false-positive rate. Powered by prodigious public proteomics data repositories, deep learning (DL) models capable of predicting peptide properties have been leveraged to compare experimental observations against predicted values, producing similarity metrics which can increase the sensitivity of peptide-spectrum matches (PSMs). This helps us achieve deeper insights into the proteome. The contribution of this work is our software MSBooster, an open-source PSM rescoring enhancer fully integrated into the end-to-end proteomics data analysis pipeline FragPipe. Following an introduction to proteomics, mass spectrometry, and machine learning in Chapter 1, each chapter builds on MSBooster. In Chapter 2, we present the algorithm behind MSBooster; showcase the unique benefits of the tool for immunopeptidomics, data-independent acquisition (DIA), single cell, and trapped ion mobility spectrometry data; and provide expected rates of improvements in peptide identification across the different tasks. We demonstrate that large human leukocyte antigen (HLA) peptide search spaces and chimeric spectra present challenges for database searching, but DL-based rescoring can improve the precision of peptide identification. Notably, we reanalyze a patient-specific melanoma tissue sample, identifying neoantigens for personalized immunotherapy. We then explore the use of specialized models in specific proteomics niches in Chapter 3, performing a benchmarking study of models available in a DL-democratizing model repository. We identify high-performing models for HLA, phosphoproteomics, tandem mass tag-labelled, and DIA Astral (Thermo Fisher Scientific) experiments. By swapping out the predictor while keeping all other up- and downstream processing steps the same in FragPipe, we isolate the prediction step to directly assess how well each model performs in the PSM rescoring context. To combat potential choice overload arising from the growing list of peptide prediction models, we present a heuristic model selection algorithm that maximizes peptide identifications. Fascinatingly, we uncover a disjunction between the loss functions used to train retention time models and actual identification improvements. To explore the “dark fragmentome” in Chapter 4, we extract additional information from MS/MS spectra beyond b- and y-ions with full spectrum prediction and showcase how that further improves HLA peptide identification. We also determine the prevalence of “unknown” peaks in full spectrum prediction. Though we found a non-negligible proportion of predicted peaks we were unable to annotate, considering said unknown peaks may actually decrease peptide identification rates. Chapter 5 summarizes our work and looks ahead to the future of PSM rescoring and DL models in bottom-up proteomics, with a focus on missing score imputation, transfer learning, and other steps in the data analysis pipeline that can benefit from peptide property prediction. Overall, MSBooster provides a practical solution for incorporating chemical knowledge into the assessment of PSM confidence. Its generalizability facilitates its use with diverse data, incorporating DL predictions into the characterization of everything from drug-binding sites to umami peptides.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

La source scientifique ouverte est momentanément indisponible.

Sujets associés

Advanced Proteomics Techniques and ApplicationsMass Spectrometry Techniques and Applications

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, ROR et la Banque mondiale, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune donnée externe enregistrée en base. Sources et limites.