Interpretable machine learning for depression symptom classification in NHANES: Performance in a curated high-confidence corpus and the full survey population
Résumé fourni par la source
Background Depressive disorders are among the most common psychiatric conditions worldwide and frequently remain undetected or misinterpreted. Scalable computational tools may help characterize depressive-symptom patterns in large health datasets, but their clinical use requires realistic validation and calibration. Objective We evaluated machine-learning models for classifying PHQ-9 depressive-symptom status in NHANES and compared performance in a curated high-confidence corpus and the full analytic population. Methods This cross-sectional analysis included 37,959 NHANES participants with valid PHQ-9 data and prespecified demographic, sleep, and dietary predictors. The primary outcome was mild-or-greater depressive symptoms, defined as PHQ-9 ≥5. A balanced high-confidence corpus was constructed using clear PHQ-9 rules, teacher-model confidence ranking, Isolation Forest filtering, and class-specific resampling. A LightGBM classifier was compared with logistic regression, random forest, gradient boosting, support vector machine, XGBoost, k-nearest neighbors, and Gaussian naïve Bayes. Performance, calibration, temporal validation, sensitivity analyses, and SHAP-based interpretability were assessed. Results The curated corpus contained 5,000 balanced observations. LightGBM achieved near-perfect curated-holdout performance, but full-population performance was moderate: accuracy 0.678, F1-score 0.503, precision 0.407, recall 0.659, and ROC-AUC 0.726. Comparator models showed similar full-population discrimination. Removing sleep predictors reduced ROC-AUC to 0.611. Full-population calibration was poor, and self-reported trouble sleeping was the dominant contributor. Conclusion Interpretable machine learning identified reproducible depressive-symptom patterns in NHANES, largely driven by sleep variables. However, modest precision and poor calibration preclude stand-alone clinical use without external validation and recalibration.