Assessing the Clinical Utility of Multimodal Large Language Models in the Diagnosis and Management of Pigmented Choroidal Lesions
Rattachement africain : us. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Purpose: To evaluate the diagnostic and treatment recommendation performance of multimodal large language models (MLLMs) in identifying and classifying retinal lesions as choroidal nevus or melanoma, as well as compare their performance with expert human graders. Methods: This retrospective cross-sectional study included 48 eyes from 47 patients diagnosed with either choroidal nevus or melanoma. Patient demographics, including age, sex, ethnicity, best-corrected visual acuity (BCVA), and symptoms, were documented. Color fundus, autofluorescence, optical coherence tomography, and B-scan images were collected. The ocular images and patient characteristics were presented to ChatGPT 4.0, Gemini Advanced 1.5 Pro, and Perplexity Pro. Responses were recorded and compared with the clinical diagnoses and treatment recommendations made by two expert human graders. Diagnostic and treatment agreement, accuracy, sensitivity, and specificity were analyzed. Results: Gemini consistently outperformed ChatGPT and Perplexity across diagnostic and treatment prompts. The highest model performance was observed for prompts requesting treatment recommendations with clinical information, where Gemini achieved the highest accuracy (0.725), followed by Perplexity (0.647) and ChatGPT (0.314). Performance was lowest for prompts requiring strict clinical criteria, with all models showing poor sensitivity. Both human graders outperformed all MLLMs in accuracy and sensitivity on most prompts (P < 0.005). Accuracy did not improve when provided demographic or clinical data, except for Gemini. Conclusions: Human graders outperform current MLLMs, which show only moderate ability to diagnose choroidal nevi or melanoma from imaging. Translational Relevance: This study highlights limitations and potential of MLLMs in aiding diagnosis and treatment of choroidal lesions.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Assessing the Clinical Utility of Multimodal Large Language Models in the Diagnosis and Management of Pigmented Choroidal Lesions
- Date Crossref
- 14/10/2025
- Éditeur
- Association for Research in Vision and Ophthalmology (ARVO)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.