Half of alcohol, drug, and self-harm presentations cannot be identified in coded emergency department data: a diagnostic accuracy study of a large language model
Rattachement africain : gb, au. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Abstract Objective To measure the accuracy of clinical coding, clinician review, and a locally deployed large language model (LLM) in identifying alcohol, drug, and self-harm involvement in emergency department (ED) attendances, and quantify prevalence. Design Two-phase diagnostic accuracy study. In a validation week, the identification strategies were assessed against a conflict-adjudicated reference standard (n=2,256); the LLM was then applied to n=105,096 annual attendances at the same site. Setting UK Type 1 Emergency Department treating patients ≥16yrs. Main outcome measures Prevalence quantification compared with the reference standard; sensitivity, specificity, and balanced accuracy of each strategy; monthly identification rates and adjusted annual prevalence. Results The reference standard identified 12.1% of attendances as involving alcohol, drugs, or self-harm (coding 6.0%; clinician 10.0%, LLM 15.6%). LLM balanced accuracy matched or outperformed clinician review in all three domains (alcohol 0.942 v 0.930, p=0.635; drug 0.959 v 0.791, p<0.001; self-harm 0.982 v 0.908, p=0.004). Coding recorded 1.07 domains per identified patient against 1.32 in the reference standard. Adjusted annual prevalence corresponded to 12,890 domain involvements per year not identifiable in coded data. Subdomain classification found at least 81.6% of self-harm attendances required medical assessment for injury or overdose before psychiatric review. Conclusions Clinical coding identified fewer than half of presentations involving alcohol, drugs, and self-harm and rarely captured co-occurring domains; under-recording was present across a full year. A locally deployed LLM generated more complete structured data from existing clinical text within NHS infrastructure, at a scale which is not feasible for manual review. What is already known on this topic Coded emergency department data underpin national surveillance, commissioning, research, and service planning, but the size and direction of their error for specific conditions have not been measured against a validated reference standard in an unselected population. What this study adds In a UK emergency department, clinical coding identified 6.0% of attendances as involving alcohol, drugs, or self-harm against an adjudicated reference-standard prevalence of 12.1%. In blinded head-to-head comparison, a locally deployed, open-source large language model achieved similar or higher balanced accuracy than clinician review. Applied to 105,096 attendances over 12 months, it identified 14.3% of attendances against 4.1% by coding, in every month of the year. How this study might affect research, practice, or policy Prevalence estimates and service planning based on coded emergency department data are likely to understate both the burden and the co-occurrence of these conditions. The majority (81.6%) of classified self-harm presentations required medical assessment for injury or overdose; only 18.4% did not, a figure directly relevant to mental health emergency-care streaming. Locally deployed open-source language models offer a scalable, governance-compatible method for identifying these presentations from existing free text.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Half of alcohol, drug, and self-harm presentations cannot be identified in coded emergency department data: a diagnostic accuracy study of a large language model
- Date Crossref
- 31/08/2026
- Éditeur
- openRxiv
- Type
- posted-content
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.