Sequential vancomycin trough and dosing interval prediction in ICU patients using general purpose large language models: a head-to-head comparison of ChatGPT and GROK AI
Résumé fourni par la source
Vancomycin monitoring remains challenging due to fluctuating renal function and ICU physiology. This study compares ChatGPT and Grok for predicting sequential vancomycin trough levels and dosing intervals across varying renal function. This retrospective study used deidentified MIMIC-IV ICU data. Admissions with three sequential vancomycin troughs and complete dosing history were included (239 admissions, 717 predictions: T1–T3). Structured clinical snapshots were submitted to both models to predict trough concentrations (mg/L) and dosing intervals (hours). Performance was assessed using MAE, RMSE, bias, ±2 mg/L and ±2-hour accuracy, sequence success, and generalized estimating equations. ChatGPT outperformed Grok in trough prediction (MAE 4.07 vs 5.08 mg/L; RMSE 5.55 vs 7.53 mg/L; ±2 mg/L accuracy 37.0% vs 29.0%), with greater advantage at T1/T2 and in renal dysfunction. Grok was superior for interval prediction (MAE 1.36 vs 1.61 hours; RMSE 2.09 vs 2.98 hours). ChatGPT achieved more successful trough sequences (≥2/3 within ±2 mg/L: 35.1% vs 23.0%), while Grok had more accurate interval sequences (all 3 within ±2 hours: 67.8% vs 59.8%). These findings represent comparative model benchmarking and require validation against pharmacist-led therapeutic drug monitoring and Bayesian AUC-guided dosing platforms before clinical use.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.