Performance of Large Language Models in Neurology Multiple‐Choice Questions
Rattachement africain : ir. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Introduction Navigating neurological disorders is complex due to overlapping symptoms and diverse diagnostic requirements. Large language models (LLMs) offer potential support for clinicians by processing vast amounts of textual data. This study evaluates the performance of four advanced LLMs—GPT‐4, GPT‐3.5, Clinical Camel, and MedALPACA—in answering neurology‐based multiple‐choice questions (MCQs). Methods The study utilized 170 MCQs from the Comprehensive Review in Clinical Neurology book. Questions were divided into the body and answer choices, stored separately in an Excel spreadsheet. The models were prompted to select the correct answers. Generative Pretrained Transformer (GPT) models were accessed via OpenAI′s API, whereas Clinical Camel and MedALPACA were downloaded from Hugging Face. Accuracy was calculated by the number of correct answers over total questions, with subgroup analysis based on subject headings. Results GPT‐4 achieved the highest accuracy at 84.7%, significantly outperforming GPT‐3.5 (58.8%), Clinical Camel (52.9%), and MedALPACA (40%). GPT‐4′s performance was significantly better ( p < 0.001). The difference between GPT‐3.5 and Clinical Camel was not significant ( p = 0.27), but both outperformed MedALPACA. Response similarity was highest between GPT‐4 and GPT‐3.5 (64.1%) and lowest between GPT‐4 and MedALPACA (41.2%). In subgroup analysis, GPT‐4 was superior across all topics, achieving full scores in six topics and its lowest in vascular neurology (40%). GPT‐3.5 performed best in eight topics, Clinical Camel in three topics, and MedALPACA was the weakest in all but two topics. Conclusion GPT‐4 demonstrated the highest accuracy in answering neurology MCQs, outperforming GPT‐3.5, Clinical Camel, and MedALPACA. Clinical Camel′s comparable performance to GPT‐3.5 highlights the potential of specialized medical models. External evaluation datasets remain essential to avoid data leakage and ensure fair benchmarking. Further research is needed to expand topic coverage, assess reasoning processes, and include human comparison to support safe and effective clinical integration of medical LLMs.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Performance of Large Language Models in Neurology Multiple‐Choice Questions
- Date Crossref
- 01/01/2026
- Éditeur
- Wiley
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.