First-Trimester Gestational Diabetes Mellitus Risk Prediction with Machine Learning Techniques: Results from the BORN2020 Cohort Study
Rattachement africain : gr. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Background: Gestational diabetes mellitus (GDM) affects many pregnancies worldwide and is associated with adverse maternal and fetal outcomes. Current screening at 24–28 weeks limits opportunities for early intervention. We evaluated whether machine learning (ML) models using first-trimester clinical and dietary data can predict GDM risk before the standard oral glucose tolerance test. Methods: We analyzed data from 797 pregnant women enrolled in the BORN2020 prospective cohort study (Thessaloniki, Greece). Ten ML algorithms were evaluated across five class-imbalance handling strategies using stratified 5-fold cross-validation, with final evaluation on an independent 20% held-out test set. Features included maternal demographics, obstetric history, lifestyle factors, and 22 dietary micronutrient intakes from the pre-pregnancy period assessed by Food Frequency Questionnaire. Results: The best-performing model (Logistic Regression without resampling) achieved an AUC-ROC of 0.664 (95% CI: 0.542–0.777), with sensitivity of 0.783 and NPV of 0.932 at the pre-specified threshold. The high NPV should be interpreted in the context of the low GDM prevalence (14.7%), as NPV is mathematically dependent on disease prevalence. A reduced nine-feature model using only routine clinical and demographic variables achieved a numerically higher AUC of 0.712 (95% CI: 0.589–0.825), with overlapping confidence intervals, indicating that detailed FFQ-derived micronutrient data did not improve prediction. Maternal age and pre-pregnancy BMI were the strongest individual predictors by SHAP analysis. No model reached the AUC >0.80 threshold for good discrimination. Substantial miscalibration was observed (slope: 0.56; intercept: −1.83), limiting use for absolute risk estimation. Conclusions: This exploratory study demonstrates that first-trimester ML models achieve modest discriminative ability for early GDM prediction, with routine clinical variables performing comparably to models incorporating detailed dietary assessment. These findings should be interpreted with caution, as no external validation cohort was available and the low events-per-variable ratio (~3.8) constrains the reliability of individual model estimates. Substantial miscalibration further limits use for absolute risk estimation. Accordingly, these models should be regarded as exploratory risk-ranking tools only and require external validation and recalibration before any clinical implementation.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- First-Trimester Gestational Diabetes Mellitus Risk Prediction with Machine Learning Techniques: Results from the BORN2020 Cohort Study
- Date Crossref
- 23/03/2026
- Éditeur
- MDPI AG
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.