Aller au contenu principal
2026 preprint

Evaluating Fairness and Generalizability of Large Language Models for Social Isolation Extraction from Electronic Health Records: Multisite Evaluation (Preprint)

0Citations signalées, ce qui n’est pas une note de qualité
0Institutions déclarées
0Pays d’affiliation déclarés

Le résumé fourni par la source

BACKGROUND Recent advancements in large language models (LLMs) have improved the identification of social isolation from clinical narratives, which vary widely in linguistic patterns and documentation practices. However, LLMs fine-tuned on a single dataset often show reduced performance when applied to different healthcare settings or clinical note types. Rigorous evaluation of cross-site generalizability and fairness is therefore essential to ensure accurate and equitable detection of social isolation across diverse populations and clinical contexts. OBJECTIVE This study aimed to evaluate a span-level fine-tuned FLAN-T5-Large model for extracting ‘social isolation’ indicators from unstructured clinical text and to assess its generalizability and fairness across diverse populations and healthcare data sources. METHODS A total of 2,967 unique annotated spans from 9,578 clinical notes across three healthcare systems were used to fine-tune an FLAN-T5-Large model using a contextualized span classification framework. A Gemma‑2-2B model was evaluated in a sensitivity analysis to assess architecture‑related performance differences. Performance was assessed using precision, recall, and macro F1. Fairness was evaluated across demographic variables, social vulnerability strata, and note types using statistical parity difference (SPD) and equal opportunity difference (EOD). RESULTS Incorporating contextual windows around annotated spans improved macro-F1 from 0.90 to 0.94 during validation. In full note evaluation across 900 manually reviewed notes, FLAN-T5-Large achieved high recall for social isolation (0.94 – 0.98) and macro F1 values ranging from 0.69 to 0.81 across sites. Fairness analysis showed generally consistent performance across age, gender, race, and social vulnerability groups, with equitable sensitivity (EOD 0.02 – 0.04) and moderate variation in positive prediction rates (SPD). Performance variability was strongly associated with documentation type, with note type driving substantially greater variability in both performance and fairness metrics than patient demographic factors. CONCLUSIONS Fine-tuned FLAN‑T5-Large demonstrated strong capability in detecting ‘social isolation’ from clinical narratives, while maintaining sensitivity parity across subgroups. The observed heterogeneity was largely driven by documentation context rather than by patient characteristics, highlighting the importance of note type-aware evaluation in clinical NLP. These findings support the use of instruction‑tuned LLMs for equitable extraction of social context information from EHR text.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Evaluating Fairness and Generalizability of Large Language Models for Social Isolation Extraction from Electronic Health Records: Multisite Evaluation (Preprint)
Date Crossref
24/04/2026
Éditeur
JMIR Publications Inc.
Type
posted-content

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les sujets associés

Machine Learning in HealthcareMental Health via WritingDigital Mental Health Interventions

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.