Diagnostic Performance of Multimodal Large Language Models in Assigning Bone-RADS Categories on CT
Rattachement africain : tr. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
The aim of the study is to assess the performance of multimodal large language models (MLLMs) in assigning Bone-RADS categories to bone lesions identified on CT images. An MSK radiologist selected one representative slice for 50 bone lesions seen on CT studies and assigned reference Bone-RADS categories using clinical records. Three raters categorized each case: an abdominal radiologist, OpenAI ChatGPT 5, and Google Gemini 2.5 Pro. Accuracy was defined as the correctly labeled Bone-RADS 1 and 4 cases and compared using McNemar test. Agreement with the reference was assessed using weighted Cohen’s κ with 95% CIs; pairwise κ differences were tested via bootstrap. Reference categories were Bone-RADS 1, n=23; 2, n=4; 3, n=0; 4, n=23. Accuracy was 84.8% (39/46) for the radiologist, 78.3% (36/46) for Gemini, and 65.2% (30/46) for ChatGPT. The radiologist outperformed ChatGPT (p=0.012); differences between the radiologist vs Gemini (p=0.604) and Gemini vs ChatGPT (p=0.360) were not significant. The radiologist achieved the highest agreement with the reference standard (κ = 0.715, 95% CI: [0.543-0.887]), followed by Gemini (κ = 0.542, 95% CI: [0.313-0.770]) and ChatGPT (κ = 0.292, 95% CI: [0.104-0.479]). Bootstrap comparisons showed that the radiologist’s κ was higher than ChatGPT’s (95% CI for difference, 0.140-0.675), while radiologist vs Gemini (−0.113-0.434) and Gemini vs ChatGPT (−0.041-0.522) were not significant. In conclusion, general-purpose MLLMs cannot yet replace trained radiologists for Bone-RADS classification, though they may still aid routine clinical practice.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Diagnostic Performance of Multimodal Large Language Models in Assigning Bone-RADS Categories on CT
- Date Crossref
- 10/02/2026
- Éditeur
- Uludag Universitesi Tip Fakultesi Dergisi
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Bursa Uludağ Üni̇versi̇tesi̇ pays non établi dans la noticeUniversité ou école supérieure
-
Necmettin Erbakan University pays non établi dans la noticeUniversité ou école supérieure
-
Bursa Uludağ University pays non établi dans la noticeUniversité ou école supérieure
Bursa Uludağ Üni̇versi̇tesi̇, Necmettin Erbakan University et Bursa Uludağ University.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.