Aller au contenu principal
2025 article

Leveraging Natural Language Processing and Generative Models in Molecular Chemistry: Property Prediction and Novel Compound Generation

4Citations signalées, ce qui n’est pas une note de qualité
5Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : us, ru. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

The accurate prediction of molecular properties is important for the rational design and the advancement of green chemistry and sustainable materials research. However, the predictive power of traditional computational chemistry methods is limited due to computational restrictions. Here, we examine an alternative approach to the accurate prediction of properties of organic compounds: natural language processing (NLP)-based molecular embedding. Using viscosity, partition coefficient (log P ), and enthalpy of vaporization as test properties through a survey of comprehensive datasets comprising 5695 data points for viscosity, 25 870 data points for log P, and 2296 data points for enthalpy of vaporization. These are important properties for the design of greener, safer, and sustainable chemical processes. Models were trained using NLP methods such as Mol2vec and fine-tuned ChemBERTa, and results were compared with traditional input featurization techniques such as Morgan fingerprints and quantum chemistry derived sigma profiles and DFT features. Among the various machine learning models, Mol2vec demonstrated superior predictive capabilities, achieving the highest correlation coefficient ( R 2 = 0.945) and lowest RMSE (0.106 mPa s) for viscosity, as well as high accuracy for log P and enthalpy of vaporization predictions. These findings establish the Mol2vec featurization technique, graph-convolutional neural networks (GCNN), and fine-tuned ChemBERTa model as powerful tools for predictive modeling of organic compounds properties, offering a significant improvement over previously used featurization techniques and opening up strategies for very-high-throughput computational screening. Finally, we integrated ML models with hybrid language-model-based generative adversarial networks (LM-GAN) to generate novel molecular sequences with desirable properties for different research applications. The ability to computationally design solvents with lower viscosity, lower log P, and lower enthalpy of vaporization offers a data-driven route to accelerating the discovery of sustainable alternatives to traditionally toxic solvents.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Leveraging Natural Language Processing and Generative Models in Molecular Chemistry: Property Prediction and Novel Compound Generation
Date Crossref
24/11/2025
Éditeur
American Chemical Society (ACS)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Oak Ridge National Laboratory pays non établi dans la notice
    Structure de recherche
  • Institute of Molecular Biology and Biophysics pays non établi dans la notice
    Structure de recherche
  • Texas A&M University – Corpus Christi pays non établi dans la notice
    Université ou école supérieure
  • Manufacturing Laboratories (United States) pays non établi dans la notice
    Entreprise
  • University of Tennessee at Knoxville pays non établi dans la notice
    Université ou école supérieure
  • Biosciences Division and Center for Molecular Biophysics pays non établi dans la notice
    Institution
  • Department of Mathematics and Statistics pays non établi dans la notice
    Institution
  • Texas A&M University-Corpus Christi pays non établi dans la notice
    Université ou école supérieure
  • Computational Sciences and Engineering Division pays non établi dans la notice
    Institution
  • Manufacturing Science Division pays non établi dans la notice
    Institution
  • Department of Biochemistry and Cellular and Molecular Biology pays non établi dans la notice
    Institution

Oak Ridge National Laboratory, Institute of Molecular Biology and Biophysics et Texas A&M University – Corpus Christi, avec 8 autres affiliations.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Machine Learning in Materials ScienceComputational Drug Discovery MethodsPhase Equilibria and Thermodynamics

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.