Evaluating the Diagnostic Accuracy and Fine-Grained Visual Reasoning of Multimodal Large Language Models in Radiology: A Prospective Comparative (Preprint)
Le résumé fourni par la source
BACKGROUND Multimodal large language models (LLMs) are rapidly expanding in radiology and digital health, but their diagnostic performance across diverse clinical scenarios and fine-grained visual reasoning tasks requires systematic evaluation. OBJECTIVE To evaluate the diagnostic accuracy and fine-grained visual interpretability of four multimodal LLMs using bilingual medical imaging questions. METHODS This prospective study evaluated bilingual medical imaging questions using zero-shot prompting on four LLMs: GPT-5.2, Gemini 3 Pro, Doubao, and Yuanbao. Subgroup analyses by language (English/Chinese), question type (image/text), imaging modality (CT, radiography, MRI, US, and nuclear medicine), and anatomical system (neurologic, respiratory, cardiovascular, digestive, genitourinary, and musculoskeletal) evaluated intermodel accuracy and intramodel consistency. Image-based visual interpretability was assessed across modality, anatomy, lesion detection, and localization. RESULTS A total of 298 questions (198 image-based, 100 text-based) were evaluated. Gemini 3 Pro achieved the highest overall accuracy (268/298, 89.9%; 95% CI 86.0%-92.9%), significantly outperforming Doubao (235/298, 78.9%; P < .001), GPT-5.2 (233/298, 78.2%; P < .001), and Yuanbao (216/298, 72.5%; P < .001). All models performed better on text-based than image-based questions (P < .05). GPT-5.2 and Yuanbao favored English questions (P < .05), while intramodel performance remained stable across modalities and anatomical systems (P > .05). Despite near-perfect macroscopic visual recognition (modality and anatomy: 99.5%-100.0%), accuracy dropped in complex tasks. Gemini 3 Pro outperformed Yuanbao and GPT-5.2 in lesion detection (175/198, 88.4%; P < .05) and outperformed all other models in lesion localization (189/198, 95.5%; P < .05). CONCLUSIONS Gemini 3 Pro demonstrated optimal comprehensive performance and cross-lingual stability. However, current multimodal LLMs face a technical bottleneck in complex visual interpretation—particularly in fine-grained lesion detection and localization—emphasizing the need to prioritize local feature perception in future model development.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Evaluating the Diagnostic Accuracy and Fine-Grained Visual Reasoning of Multimodal Large Language Models in Radiology: A Prospective Comparative (Preprint)
- Date Crossref
- 20/04/2026
- Éditeur
- JMIR Publications Inc.
- Type
- posted-content
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.