A Systematic Performance Evaluation of Three Large Language Models in Answering Questions on moderate Hyperthermia
Rattachement africain : ch, de, nl, us, hu, pl, se. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
ABSTRACT Background Large Language Models (LLMs) have demonstrated expert-level performance across many medical domains, suggesting potential utility in clinical practice. However, their reliability in the highly specialized domain of moderate hyperthermia (HT) remains unknown. We therefore evaluated the performance of three modern LLMs in answering HT-related questions. Methods We conducted an evaluation study by posing 40 open-ended questions—22 clinical and 18 physics-related—to three modern LLMs (DeepSeek-V3, Llama-3.3-70B-Instruct, and GPT-4o). Responses were blinded, randomized, and evaluated by 19 international experts with either a clinical or physics background for quality (5-point Likert scale: 1=very bad, 2=bad, 3=acceptable, 4=good to 5=very good) and for potential harmfulness in clinical decision-making. Results A total of 1144 quality evaluation responses were collected. Overall reported mean quality scores were similar across models, with DeepSeek scoring 3.26, Llama 3.18, and GPT-4o 3.07, corresponding to an “acceptable” rating. Across expert evaluations, responses were considered potentially harmful in 17.8% of cases for DeepSeek, 19.3% for Llama, and 15.3% for GPT-4o. Notably, despite “acceptable” mean scores, approximately 25% of responses were rated “bad” to “very bad,” and potentially harmful answers occurred in ∼15–19% of evaluations, indicating a non-trivial risk if used without domain expertise. Conclusion Our findings indicate that the performance of LLMs in HT in versions available at the time of investigation is only partially satisfactory. The proportion of poor-quality responses is too high and may lead non-domain experts to misinterpret the available clinical evidence and draw inappropriate clinical conclusions.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- A Systematic Performance Evaluation of Three Large Language Models in Answering Questions on moderate Hyperthermia
- Date Crossref
- 26/03/2026
- Éditeur
- openRxiv
- Type
- posted-content
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
University of Bern pays non établi dans la noticeUniversité ou école supérieure
-
University Hospital of Bern Department of Internal Medicine III pays non établi dans la noticeÉtablissement de santé
-
LMU Klinikum pays non établi dans la noticeÉtablissement de santé
-
Ludwig-Maximilians-Universität München pays non établi dans la noticeUniversité ou école supérieure
-
Amsterdam Neuroscience pays non établi dans la noticeStructure de recherche
-
University of Amsterdam Amsterdam UMC pays non établi dans la noticeUniversité ou école supérieure
-
Erasmus MC Cancer Institute pays non établi dans la noticeÉtablissement de santé
-
Charité - Universitätsmedizin Berlin pays non établi dans la noticeÉtablissement de santé
-
Immunologie-Zentrum Zürich pays non établi dans la noticeÉtablissement de santé
-
Westchester Medical Center Department of Radiation Medicine pays non établi dans la noticeÉtablissement de santé
-
Kantonsspital Aarau pays non établi dans la noticeÉtablissement de santé
-
University of Maryland Department of Radiation Oncology pays non établi dans la noticeUniversité ou école supérieure
University of Bern, Department of Internal Medicine III — University Hospital of Bern et LMU Klinikum, avec 9 autres affiliations.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.