Aller au contenu principal
Accès ouvert déclaré 2025 article

Benchmarking GPT-5 in radiation oncology: measurable gains, but persistent need for expert oversight

3Citations signalées, ce qui n’est pas une note de qualité
5Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : de. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Introduction Large language models (LLM) have shown great potential in clinical decision support and medical education. GPT-5 is a novel LLM system that has been specifically marketed towards oncology use. This study comprehensively benchmarks GPT-5 for the field of radiation oncology. Methods Performance was assessed using two complementary benchmarks: (i) the American College of Radiology Radiation Oncology In-Training Examination (TXIT, 2021), comprising 300 multiple-choice items, and (ii) a curated set of 60 authentic radiation oncologic vignettes representing diverse disease sites and treatment indications. For the vignette evaluation, GPT-5 was instructed to generate structured therapeutic plans and concise two-line summaries. Four board-certified radiation oncologists independently rated outputs for correctness, comprehensiveness, and hallucinations. Inter-rater reliability was quantified using Fleiss’ κ . GPT-5–14 results were compared to published GPT-3.5 and GPT-4 baselines. Results On the TXIT benchmark, GPT-5 achieved a mean accuracy of 92.8%, outperforming GPT-4 (78.8%) and GPT-3.5 (62.1%). Domain-specific gains were most pronounced in dose specification and diagnosis. In the vignette evaluation, GPT-5’s treatment recommendations were rated highly for correctness (mean 3.24/4, 95% CI: 3.11–3.38) and comprehensiveness (3.59/4, 95% CI: 3.49–3.69). Hallucinations were rare, flagged in 10.0% of all individual reviewer assessments (24 of 240), and no patient case reached majority consensus for their presence. Inter-rater agreement was low (Fleiss’ κ 0.083 for correctness), reflecting inherent variability in clinical judgment. Errors clustered in complex scenarios requiring precise trial knowledge or detailed clinical adaptation. Discussion GPT-5 clearly outperformed prior model variants on the radiation oncology multiple-choice benchmark. Although GPT-5 exhibited favorable performance in generating real-world radiation oncology treatment recommendations, correctness ratings indicate room for further improvement. While hallucinations were infrequent, the presence of substantive errors underscores that GPT-5-generated recommendations require rigorous expert oversight before clinical implementation. In addition, considerable inter-rater variability highlights the challenge of achieving consistent expert evaluation.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Benchmarking GPT-5 in radiation oncology: measurable gains, but persistent need for expert oversight
Date Crossref
11/12/2025
Éditeur
Frontiers Media SA
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Radiomics and Machine Learning in Medical ImagingManagement of metastatic bone diseaseMedical Imaging Techniques and Applications

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.