Aller au contenu principal
Accès ouvert déclaré 2026 preprint

Evaluating the Diagnostic Accuracy and Fine-Grained Visual Reasoning of Multimodal Large Language Models in Radiology: A Prospective Comparative (Preprint)

0Citations signalées, ce qui n’est pas une note de qualité
0Institutions déclarées
0Pays d’affiliation déclarés

Le résumé fourni par la source

BACKGROUND Multimodal large language models (LLMs) are rapidly expanding in radiology and digital health, but their diagnostic performance across diverse clinical scenarios and fine-grained visual reasoning tasks requires systematic evaluation. OBJECTIVE To evaluate the diagnostic accuracy and fine-grained visual interpretability of four multimodal LLMs using bilingual medical imaging questions. METHODS This prospective study evaluated bilingual medical imaging questions using zero-shot prompting on four LLMs: GPT-5.2, Gemini 3 Pro, Doubao, and Yuanbao. Subgroup analyses by language (English/Chinese), question type (image/text), imaging modality (CT, radiography, MRI, US, and nuclear medicine), and anatomical system (neurologic, respiratory, cardiovascular, digestive, genitourinary, and musculoskeletal) evaluated intermodel accuracy and intramodel consistency. Image-based visual interpretability was assessed across modality, anatomy, lesion detection, and localization. RESULTS A total of 298 questions (198 image-based, 100 text-based) were evaluated. Gemini 3 Pro achieved the highest overall accuracy (268/298, 89.9%; 95% CI 86.0%-92.9%), significantly outperforming Doubao (235/298, 78.9%; P < .001), GPT-5.2 (233/298, 78.2%; P < .001), and Yuanbao (216/298, 72.5%; P < .001). All models performed better on text-based than image-based questions (P < .05). GPT-5.2 and Yuanbao favored English questions (P < .05), while intramodel performance remained stable across modalities and anatomical systems (P > .05). Despite near-perfect macroscopic visual recognition (modality and anatomy: 99.5%-100.0%), accuracy dropped in complex tasks. Gemini 3 Pro outperformed Yuanbao and GPT-5.2 in lesion detection (175/198, 88.4%; P < .05) and outperformed all other models in lesion localization (189/198, 95.5%; P < .05). CONCLUSIONS Gemini 3 Pro demonstrated optimal comprehensive performance and cross-lingual stability. However, current multimodal LLMs face a technical bottleneck in complex visual interpretation—particularly in fine-grained lesion detection and localization—emphasizing the need to prioritize local feature perception in future model development.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Evaluating the Diagnostic Accuracy and Fine-Grained Visual Reasoning of Multimodal Large Language Models in Radiology: A Prospective Comparative (Preprint)
Date Crossref
20/04/2026
Éditeur
JMIR Publications Inc.
Type
posted-content

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les sujets associés

Artificial Intelligence in Healthcare and EducationRadiomics and Machine Learning in Medical ImagingCOVID-19 diagnosis using AI

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.