How to build machine learning models able to extrapolate from standard to modified peptides
Résumé fourni par la source
Bioactive peptides are an important class of natural products with great functional versatility. Chemical modifications can improve their pharmacology, yet their structural diversity presents unique challenges for computational modeling. Furthermore, data for standard peptides (composed of the 20 canonical amino acids) is more abundant than for modified ones. Thus, we set out to identify whether predictive models fitted to standard data are reliable when applied to modified peptides. To do this, we first considered two critical aspects of the modeling problem, namely, choice of similarity function for guiding dataset partitioning and choice of molecular representation. Similarity-based dataset partitioning is an evaluation technique that divides the dataset into train and test subsets, such that the molecules in the test set are different from those used to fit the model. Scientific contribution. We demonstrate, across four peptide function prediction tasks, that chemical fingerprint-based similarity measures outperform traditional sequence alignment-based metrics for partitioning standard peptide datasets, challenging conventional practice. We have also found that the choice of representation does not have a substantial impact in the standard to standard interpolation modelling scenario, but for modified to modified interpolation chemical language models are the best option. Despite a 50\% drop in performance, chemical fingerprints and a chemical language model (ChemBERTa-2) are the best choices in the more challenging extrapolation scenario (standard to modified). All code and data necessary for reproducing the experiments are available in Github (https://github.com/IBM/PeptideGeneralizationBenchmarks).
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- How to build machine learning models able to extrapolate from standard to modified peptides
- Date Crossref
- 23/10/2025
- Éditeur
- American Chemical Society (ACS)
- Type
- posted-content
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.