Aller au contenu principal
Accès ouvert déclaré 2026 preprint

A Systematic Performance Evaluation of Three Large Language Models in Answering Questions on moderate Hyperthermia

0Citations signalées, ce qui n’est pas une note de qualité
18Institutions déclarées
7Pays d’affiliation déclarés

Rattachement africain : ch, de, nl, us, hu, pl, se. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

ABSTRACT Background Large Language Models (LLMs) have demonstrated expert-level performance across many medical domains, suggesting potential utility in clinical practice. However, their reliability in the highly specialized domain of moderate hyperthermia (HT) remains unknown. We therefore evaluated the performance of three modern LLMs in answering HT-related questions. Methods We conducted an evaluation study by posing 40 open-ended questions—22 clinical and 18 physics-related—to three modern LLMs (DeepSeek-V3, Llama-3.3-70B-Instruct, and GPT-4o). Responses were blinded, randomized, and evaluated by 19 international experts with either a clinical or physics background for quality (5-point Likert scale: 1=very bad, 2=bad, 3=acceptable, 4=good to 5=very good) and for potential harmfulness in clinical decision-making. Results A total of 1144 quality evaluation responses were collected. Overall reported mean quality scores were similar across models, with DeepSeek scoring 3.26, Llama 3.18, and GPT-4o 3.07, corresponding to an “acceptable” rating. Across expert evaluations, responses were considered potentially harmful in 17.8% of cases for DeepSeek, 19.3% for Llama, and 15.3% for GPT-4o. Notably, despite “acceptable” mean scores, approximately 25% of responses were rated “bad” to “very bad,” and potentially harmful answers occurred in ∼15–19% of evaluations, indicating a non-trivial risk if used without domain expertise. Conclusion Our findings indicate that the performance of LLMs in HT in versions available at the time of investigation is only partially satisfactory. The proportion of poor-quality responses is too high and may lead non-domain experts to misinterpret the available clinical evidence and draw inappropriate clinical conclusions.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
A Systematic Performance Evaluation of Three Large Language Models in Answering Questions on moderate Hyperthermia
Date Crossref
26/03/2026
Éditeur
openRxiv
Type
posted-content

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • University of Bern pays non établi dans la notice
    Université ou école supérieure
  • University Hospital of Bern Department of Internal Medicine III pays non établi dans la notice
    Établissement de santé
  • LMU Klinikum pays non établi dans la notice
    Établissement de santé
  • Ludwig-Maximilians-Universität München pays non établi dans la notice
    Université ou école supérieure
  • Amsterdam Neuroscience pays non établi dans la notice
    Structure de recherche
  • University of Amsterdam Amsterdam UMC pays non établi dans la notice
    Université ou école supérieure
  • Erasmus MC Cancer Institute pays non établi dans la notice
    Établissement de santé
  • Charité - Universitätsmedizin Berlin pays non établi dans la notice
    Établissement de santé
  • Immunologie-Zentrum Zürich pays non établi dans la notice
    Établissement de santé
  • Westchester Medical Center Department of Radiation Medicine pays non établi dans la notice
    Établissement de santé
  • Kantonsspital Aarau pays non établi dans la notice
    Établissement de santé
  • University of Maryland Department of Radiation Oncology pays non établi dans la notice
    Université ou école supérieure

University of Bern, Department of Internal Medicine III — University Hospital of Bern et LMU Klinikum, avec 9 autres affiliations.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Artificial Intelligence in Healthcare and EducationGenomics and Rare DiseasesText Readability and Simplification

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.