Aller au contenu principal
Accès ouvert déclaré 2026 article

A Leakage-Aware Benchmark Study of Machine Learning Models for Deep Eutectic Solvent Property Prediction

0Citations signalées, ce qui n’est pas une note de qualité
2Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : es, pt. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Predicting the physicochemical properties of deep eutectic solvents (DESs) remains challenging due to the large combinatorial design space and the complex, composition- and temperature-dependent interactions governing their behavior. While machine learning (ML) has been widely applied to DES property prediction, reported performance is often sensitive to data set structure, feature representation, and validation design, raising questions about the reliability and transferability of existing models. In this work, we present a systematic evaluation of descriptor-based ML models for DES property prediction using a curated and high-confidence data set spanning 2003-2026. A unified feature representation combining molecular descriptors, molar composition, and temperature is employed, and model performance is assessed under a hierarchy of validation protocols designed to control for data leakage and progressively increase extrapolation difficulty. The results show that predictive performance is strongly dependent on the validation design. Density and refractive index exhibit relatively stable behavior, while surface tension shows moderate predictability. In contrast, electrical conductivity and viscosity display a strong dependence on temperature and limited contribution from descriptor-based features. Under extrapolative validation, performance for these properties deteriorates substantially, indicating limited transferability. Additional analyses based on feature-space distance and similarity-based baselines indicate that prediction errors are strongly associated with training-domain proximity and that local similarity in the current feature space is insufficient for reliable prediction outside observed data regions. These findings suggest that the primary limitation arises from the representational capacity of static descriptor-based features rather than model choice alone. Overall, this study provides a structured and reproducible assessment of the conditions under which descriptor-based ML models can be expected to succeed or fail in DES systems and highlights the importance of rigorous, leakage-aware evaluation in data-driven chemical modeling.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
A Leakage-Aware Benchmark Study of Machine Learning Models for Deep Eutectic Solvent Property Prediction
Date Crossref
07/08/2026
Éditeur
American Chemical Society (ACS)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Machine Learning in Materials ScienceIonic liquids properties and applicationsCrystallography and molecular interactions

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.