ClinicalBench: Can LLMs Beat Traditional ML Models in Clinical Prediction?
Rattachement africain : us, gb. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Large Language Models (LLMs) hold great promise to revolutionize current clinical systems for their superior capacities on medical text processing tasks and medical licensing exams. Meanwhile, traditional ML models such as SVM and XGBoost have still been mainly adopted in clinical prediction tasks. An emerging question is: Can LLMs beat traditional ML models in clinical prediction? Thus, we build a new benchmark ClinicalBench to comprehensively study the clinical predictive modeling capacities of both general-purpose and medical LLMs, and compare them with traditional ML models. ClinicalBench embraces three common clinical prediction tasks, two databases, 14 general-purpose LLMs, 8 medical LLMs, and 11 traditional ML models. Through extensive empirical investigation, we discover that both general-purpose and medical LLMs, even with different model scales, diverse prompting or fine-tuning strategies, still cannot beat traditional ML models in clinical prediction yet, shedding light on their potential deficiency in clinical reasoning and decision-making. We call for caution when practitioners adopt LLMs in clinical applications. ClinicalBench can be utilized to bridge the gap between LLMs' development for healthcare and real-world clinical practice. Extended version with a more detailed appendix: https://arxiv.org/abs/2411.06469 Project website: https://clinicalbench.github.io/ The code: https://github.com/canyuchen/ClinicalBench
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- ClinicalBench: Can LLMs Beat Traditional ML Models in Clinical Prediction?
- Date Crossref
- 08/08/2026
- Éditeur
- ACM
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Northwestern University Department of Computer Science pays non établi dans la noticeUniversité ou école supérieure
-
The University of Texas at Austin Department of Computer Science pays non établi dans la noticeUniversité ou école supérieure
-
Boston Children's Hospital pays non établi dans la noticeÉtablissement de santé
-
Imperial College London Department of Computer Science pays non établi dans la noticeUniversité ou école supérieure
-
The Ohio State University Department of Computer Science pays non établi dans la noticeUniversité ou école supérieure
-
Harvard University pays non établi dans la noticeUniversité ou école supérieure
-
Massachusetts General Hospital pays non établi dans la noticeÉtablissement de santé
-
University of Minnesota Department of Surgery pays non établi dans la noticeUniversité ou école supérieure
-
Cornell University Department of Population Health Sciences pays non établi dans la noticeUniversité ou école supérieure
-
Emory University Department of Computer Science pays non établi dans la noticeUniversité ou école supérieure
-
Feinberg School of Medicine Department of Preventive Medicine pays non établi dans la noticeUniversité ou école supérieure
Department of Computer Science — Northwestern University, Department of Computer Science — The University of Texas at Austin et Boston Children's Hospital, avec 8 autres affiliations.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.