Large Language Model Performance in UK Advice & Guidance: A Pilot Study in Neurology
Rattachement africain : gb. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Abstract Background Large language models (LLMs) demonstrate strong performance in controlled medical environments such as multiple choice exams, but their utility in real-world clinical workflows remains unproven. The NHS Advice & Guidance (A&G) service, where Primary Care clinicians can submit text-based queries to specialists, provides an environment for evaluating the clinical performance of LLMs as a specialist. Methods We compared responses from MedGemma 4B-IT, an open-weight model deployed locally on hospital infrastructure, against specialist neurologist responses across 50 adult neurology A&G cases from University College London Hospital. Two neurologists and two GPs rated 80 blinded and 20 unblinded responses for outcome, safety, efficacy, and feasibility using standardised criteria; outcome was a binary correct/incorrect, while other domains were scored 1-5. Inter-rater reliability was assessed using intraclass correlation coefficients. Results Although there were no statistically significant differences between blinded specialist neurologists and LLM responses across any domain (outcome: 84% vs 82%, p=0.67; safety: 3.98 vs 4.02, p=0.85; efficacy: 4.06 vs 3.98, p=0.61; feasibility: 4.39 vs 4.20, p=0.45), 10% of LLM responses received concerning scores (≤2 average score) compared to 0% of human responses, indicating potentially clinically important tail risk. Furthermore, unblinded results showed a preference for human responses, with human ratings being preferred across all domains. Only 51% of binary outcomes had unanimous agreement and inter-rater agreement was moderate across other domains (ICC 0.50-0.52). Conclusions In this pilot study, aggregate scores between blinded human and LLM responses were similar, and no statistically significant differences were detected in this exploratory sample. However, aggregate metrics masked clinically important edge-case failures in LLM responses. Pronounced inter-rater variability and the potential impact of LLM/human syntax on blinded rater judgements highlight the challenges in establishing robust evaluation frameworks for clinical LLM deployment
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé, mais le titre doit être comparé manuellement.
- Titre Crossref
- Large Language Model Performance in UK Advice & Guidance: A Pilot Study in Neurology
- Date Crossref
- 18/05/2026
- Éditeur
- openRxiv
- Type
- posted-content
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
University College London Hospitals NHS Foundation Trust Institute of Health Informatics pays non établi dans la noticeÉtablissement de santé
-
University College London pays non établi dans la noticeUniversité ou école supérieure
-
Hillingdon Hospitals NHS Foundation Trust pays non établi dans la noticeÉtablissement de santé
-
Royal Free London NHS Foundation Trust pays non établi dans la noticeÉtablissement de santé
-
King's College London Department of Biostatistics and Health Informatics pays non établi dans la noticeUniversité ou école supérieure
-
Independent Researcher pays non établi dans la noticeInstitution
Institute of Health Informatics — University College London Hospitals NHS Foundation Trust, University College London et Hillingdon Hospitals NHS Foundation Trust, avec 3 autres affiliations.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.