1057 Performance of Grok-2, ChatGPT-4o, and Gemini 1.5 Flash Large Language Models (LLMs) in Detecting and Classifying Intracranial Hemorrhages
Le résumé fourni par la source
INTRODUCTION: Rapid interpretation of CT head scans for intracranial bleeding is critical for prioritizing neurosurgical intervention. Large language models (LLMs) have recently advanced in image analysis, with Grok-2, introduced on the social media platform X, claiming high accuracy in medical imaging interpretation. If validated, such models could serve a valuable triage function in low-resource settings and non-trauma hospitals. METHODS: Non-contrast, axial CT head scans were sourced from the RSNA 2019 database. A random sample of 400 scans (200 normal and 200 hemorrhage cases) was selected, including 40 each of epidural, intraparenchymal, intraventricular, subarachnoid, and subdural hemorrhages. Zero-shot prompting was used for each LLM to assess (1) presence of intracranial hemorrhage and (2) type of hemorrhage. A blinded medical student independently reviewed all images for comparison. McNemar’s test was applied for paired accuracy comparisons, and Cohen’s kappa assessed inter-rater agreement. RESULTS: There was no significant difference in hemorrhage detection accuracy among the LLMs, with accuracies ranging from 59.3% to 61.0%. The medical student outperformed all models in both detection accuracy and specificity. Classification performance varied, with the student significantly outperforming LLMs in detecting epidural and intraparenchymal hemorrhages. Cohen’s kappa indicated only slight agreement (κ < 0.2) between the student and each LLM, suggesting reliance on different diagnostic cues. CONCLUSIONS: Current general-purpose LLMs demonstrate moderate but inconsistent ability to detect and classify intracranial hemorrhages, underperforming compared to a human medical student. While exhibiting distinct diagnostic patterns, none of the LLMs matched human specificity or accuracy. Refinement of task-specific systems may be required to enhance clinical applicability in neuroimaging.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- 1057 Performance of Grok-2, ChatGPT-4o, and Gemini 1.5 Flash Large Language Models (LLMs) in Detecting and Classifying Intracranial Hemorrhages
- Date Crossref
- 01/04/2026
- Éditeur
- Ovid Technologies (Wolters Kluwer Health)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.