Evaluating Fairness and Generalizability of Large Language Models for Social Isolation Extraction from Electronic Health Records: Multisite Evaluation (Preprint)
Le résumé fourni par la source
BACKGROUND Recent advancements in large language models (LLMs) have improved the identification of social isolation from clinical narratives, which vary widely in linguistic patterns and documentation practices. However, LLMs fine-tuned on a single dataset often show reduced performance when applied to different healthcare settings or clinical note types. Rigorous evaluation of cross-site generalizability and fairness is therefore essential to ensure accurate and equitable detection of social isolation across diverse populations and clinical contexts. OBJECTIVE This study aimed to evaluate a span-level fine-tuned FLAN-T5-Large model for extracting ‘social isolation’ indicators from unstructured clinical text and to assess its generalizability and fairness across diverse populations and healthcare data sources. METHODS A total of 2,967 unique annotated spans from 9,578 clinical notes across three healthcare systems were used to fine-tune an FLAN-T5-Large model using a contextualized span classification framework. A Gemma‑2-2B model was evaluated in a sensitivity analysis to assess architecture‑related performance differences. Performance was assessed using precision, recall, and macro F1. Fairness was evaluated across demographic variables, social vulnerability strata, and note types using statistical parity difference (SPD) and equal opportunity difference (EOD). RESULTS Incorporating contextual windows around annotated spans improved macro-F1 from 0.90 to 0.94 during validation. In full note evaluation across 900 manually reviewed notes, FLAN-T5-Large achieved high recall for social isolation (0.94 – 0.98) and macro F1 values ranging from 0.69 to 0.81 across sites. Fairness analysis showed generally consistent performance across age, gender, race, and social vulnerability groups, with equitable sensitivity (EOD 0.02 – 0.04) and moderate variation in positive prediction rates (SPD). Performance variability was strongly associated with documentation type, with note type driving substantially greater variability in both performance and fairness metrics than patient demographic factors. CONCLUSIONS Fine-tuned FLAN‑T5-Large demonstrated strong capability in detecting ‘social isolation’ from clinical narratives, while maintaining sensitivity parity across subgroups. The observed heterogeneity was largely driven by documentation context rather than by patient characteristics, highlighting the importance of note type-aware evaluation in clinical NLP. These findings support the use of instruction‑tuned LLMs for equitable extraction of social context information from EHR text.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Evaluating Fairness and Generalizability of Large Language Models for Social Isolation Extraction from Electronic Health Records: Multisite Evaluation (Preprint)
- Date Crossref
- 24/04/2026
- Éditeur
- JMIR Publications Inc.
- Type
- posted-content
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.