Development of an explainable AI model for predicting 3-month functional recovery in middle-to older-age trauma patients visiting emergency department
Résumé fourni par la source
Middle-to-older-age trauma patients may experience persistent functional dependence after injury. Explainable machine-learning methods may reveal clinically coherent patterns across frailty, pre-injury function, injury characteristics, and subsequent care. We explored the feasibility of an explainable machine-learning framework for 3-month functional recovery and the hypotheses generated by its apparent performance and explanation outputs. We analyzed 588 patients aged ≥ 45 years from a prospective single-center trauma cohort; 73 (12.4%) had poor functional recovery, defined as a 3-month Barthel Index ≤ 60. Seventy candidate variables and minimum-redundancy maximum-relevance-selected subsets of 30, 12, and 7 variables were evaluated using an automated machine-learning framework. Safe-level synthetic minority oversampling and feature selection were applied to the full cohort before five-fold cross-validation, and no independent hold-out set was available. Patients with poor recovery were older than those with favorable recovery (82.4 ± 8.3 vs. 70.6 ± 12.5 years) and had lower pre-injury Barthel Index scores (69.0 ± 27.7 vs. 97.1 ± 8.7). Within the original exploratory workflow, the 30-variable Fine KNN configuration yielded an apparent AUROC of 0.94 and apparent sensitivity of 0.90; the 70-variable configuration yielded an apparent AUROC of 0.94 and sensitivity of 0.86, whereas the 7-variable configurations showed lower apparent sensitivity (0.63–0.64). Collectively, the MRMR, SHAP, and LIME outputs highlighted pre-injury function, age, living arrangement, and injury-related variables. Taken together, these patterns support the hypotheses that baseline functional reserve and multidomain injury information may be important for 3-month recovery, and that very aggressive feature reduction may discard relevant signal. Because full-cohort preprocessing, oversampling, feature selection, and model comparison preceded cross-validation, the performance estimates and feature rankings are optimistically biased and require confirmation. This pilot analysis provides a clinically coherent proof of concept for explainable machine learning in post-traumatic functional recovery and identifies testable hypotheses for future model development. A frailty- and function-informed multidomain approach may be more informative than a highly reduced feature set. Leakage-free fold-contained redevelopment, calibration and decision-curve assessment, and independent multicenter validation are needed before clinical use.