Performance of General-Purpose Vision Language Models and Ophthalmology Foundation Models in Glaucoma Detection and Function Prediction
Rattachement africain : us. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Purpose: To evaluate the performance of vision-language models (VLMs), in glaucoma detection and visual field (VF) mean deviation (MD) prediction tasks using optical coherence tomography (OCT) images. Methods: A total of 27,610 SPECTRALIS OCT images from 1025 participants (1690 eyes), collected between 2008 and 2021 as part of the Diagnostic Innovations in Glaucoma Study (DIGS) and the African Descent and Glaucoma Evaluation Study (ADAGES), were included. Vision components of LLaVA and PaliGemma, as well as RETFound and ResNet-50 models, were fine-tuned for glaucoma classification and VF MD prediction. Models were trained using OCT circle scans centered on the optic nerve head. Three training configurations were compared. Performance was evaluated using area under the receiver operating characteristic curve (AUC), mean absolute error (MAE), and related metrics. Results: The LLaVA model, when both vision encoder and multi-layer projector were fine-tuned, achieved the best performance with an AUC of 0.92 (95% confidence interval [CI], 0.86-0.95) for glaucoma classification and an MAE of 1.79 dB (95% CI, 1.55-2.00) for VF MD prediction. RETFound and PaliGemma also performed well, with AUCs of 0.91 and 0.90 and MAEs of 1.87 dB and 1.84 dB, respectively. Models with frozen vision encoders showed reduced accuracy. Stratified analysis showed better glaucoma classification in older individuals and moderate-to-advanced cases. VF MD prediction was more accurate in younger individuals, with higher errors in advanced glaucoma. Conclusions: Fine-tuned VLMs demonstrated high performance in glaucoma detection and VF MD prediction, matching or exceeding specialized foundation models and traditional convolutional neural network (CNN)-based methods. Translational Relevance: This study highlights the potential of general-purpose AI models to be adapted for glaucoma care, enabling scalable decision support from OCT imaging.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Performance of General-Purpose Vision Language Models and Ophthalmology Foundation Models in Glaucoma Detection and Function Prediction
- Date Crossref
- 19/11/2025
- Éditeur
- Association for Research in Vision and Ophthalmology (ARVO)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.