Assessing the performance of artificial intelligence in detecting electrographic status epilepticus when human experts often disagree
Résumé fourni par la source
Objective Generating reference standards to train and evaluate the accuracy of artificial intelligence (AI) algorithms poses a significant challenge. Particularly when interpreting complex signals like electroencephalography (EEG), where interrater variability is considerable. We aimed to characterize the impact of interrater variability when evaluating the performance of an AI algorithm for detecting electrographic status epilepticus in point-of-care (POC) limited-montage EEG. Methods We analyzed 604 EEGs collected using a POC EEG system (Ceribell Inc.). Each EEG was independently reviewed by 5–7 blinded experts, who annotated seizures and completed standardized assessments. The EEGs were later analyzed by the Clarity AI algorithm (version 7), trained on separate EEG dataset. Interrater agreement was assessed using Gwet’s AC1. Sensitivity and specificity for detecting electrographic status epilepticus (ESE) were estimated for the reviewers and the AI algorithm. Multiple evaluation schemes were employed, including simple majority consensus of the full group (group majority), simple majority using leave-one-out analysis, and 2-of-3 majority with bootstrap sampling. Results Of 604 POC EEGs, the group majority identified 8 cases (1.3%) meeting ACNS criteria for ESE and 14 (2.3%) with seizures, the rest were classified as normal/slowing (80.1%), highly epileptiform patterns (6.6%), or other findings (4.3%). Twentynine cases lacked consensus interpretation. Interrater agreement among reviewers was 0.67–0.68. Compared to the group majority, the AI algorithm showed higher sensitivity (median 100%) with lower specificity (93.5%) than individual reviewers (median sensitivity 60%, specificity 98.7%). Using all possible 3-reviewer combinations to define majority agreement, the number of EEGs identified as ESE varied widely (4–18 cases, median = 10). The AI algorithm consistently achieved significantly higher sensitivity than external human reviewers (71.4% vs. 50%, p < 0.001). However, the AI’s specificity, while still high (median = 93.9%), was slightly lower than that of human reviewers (median = 98.4%, p < 0.001), though the AI’s specificity had consistently narrower spread. Discussion This study highlights the challenges of defining the correct answer for EEG interpretation, especially for ESE, and its consequences for evaluating seizure detection AI tools. Future work should explore how AI assistance impacts human interpretation, particularly in reducing interrater variability across a broader range of EEG patterns.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Assessing the performance of artificial intelligence in detecting electrographic status epilepticus when human experts often disagree
- Date Crossref
- 11/08/2026
- Éditeur
- Frontiers Media SA
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.