Performance of AI-based diabetic retinopathy screening is highly dependent on evaluation setting: a five-year, multi-framework study
Rattachement africain : fr, fi. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Abstract Despite near-perfect performance reported on benchmark datasets, the real-world behavior of artificial intelligence (AI) systems for diabetic retinopathy (DR) screening remains insufficiently characterized. Here, we report a multi-framework evaluation of OphtAI, a CE-marked AI system for automated detection of referable DR and diabetic macular edema from color fundus photographs. The system was initially validated on the Messidor-2 benchmark dataset and subsequently assessed across three independent evaluation settings: a large-scale masked comparative study (US Veterans Affairs), an external validation using handheld fundus photography in Finland, and an open, large-scale comparative evaluation within the UK National Health Service. Across these evaluation frameworks, substantial variability in observed sensitivity and specificity was found. These variations occurred across settings differing in imaging devices, population characteristics, referral definitions, and handling of ungradable images. In masked settings, interpretation of comparative performance was limited by the lack of system-level attribution, while open evaluations remained sensitive to protocol design. Taken together, these results show that performance in AI-based DR screening is not a fixed property of a system, but an emergent property of the evaluation framework. This work calls for context-aware validation strategies and more standardized evaluation protocols for clinical AI systems.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Performance of AI-based diabetic retinopathy screening is highly dependent on evaluation setting: a five-year, multi-framework study
- Date Crossref
- 10/09/2026
- Éditeur
- Springer Science and Business Media LLC
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.