Aller au contenu principal
Accès ouvert déclaré 2026 article

Transcript-based DASH scoring with personalized large language models: a mixed-methods concordance study and measurement considerations

0Citations signalées — pas une note de qualité
4Institutions déclarées
1Pays d’affiliation déclarés

Résumé fourni par la source

Abstract Background Debriefing is central to healthcare simulation, yet quality assessment relies on expert raters and substantial faculty time. Large language models (LLMs) may support transcript-based quality review using validated instruments such as the Debriefing Assessment for Simulation in Healthcare (DASH). Still, evidence is needed on agreement patterns and the constraints of transcript-only assessment. This study compared concordance between human experts and personalized ChatGPT models for DASH scoring of debriefing transcripts. Methods A prospective, multicentric pilot study with a mixed-methods design analyzed 33 high-fidelity simulation debriefings. Transcripts were evaluated by human experts ( n = 150 evaluations) and personalized GPT-4.0 and GPT-4.5 models ( n = 345 evaluations total) using DASH items 2–6 (score range 5–35). An adaptive design selected the best-performing AI model based on internal consistency (Cronbach’s α ≥ 0.7) and inter-evaluator concordance (ICC ≥ 0.5). Quantitative analyses included continuous and categorical concordance measures, and qualitative analysis used comparative content analysis. Results GPT-4.5 demonstrated superior performance (α = 0.870) compared with GPT-4.0 (α = 0.427) and was selected for the main analysis. Continuous analyses showed moderate concordance with human experts (ICC range 0.290–0.560; median 0.455). Categorical analyses showed high agreement (95–100%; PABAK 0.900–1.000) in a dataset in which ratings were predominantly classified as “Good.” Qualitatively, both experts and GPT-4.5 frequently identified similar strengths, while experts more often provided specific, critical improvement-oriented feedback, and the model tended toward a predominantly positive, descriptive tone. Conclusions In transcript-only scoring of DASH items 2–6, a personalized GPT-4.5 model demonstrated strong internal consistency and moderate concordance with expert ratings on continuous scales, while coarse categorical classifications yielded high agreement in a predominantly high-scoring sample. These findings support a human-supervised role for LLMs in structuring transcript-based review and potentially triaging debriefings for faculty development workflows, rather than substituting for expert assessment. Future work should test performance in more heterogeneous quality distributions and with multimodal observation beyond transcripts.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Transcript-based DASH scoring with personalized large language models: a mixed-methods concordance study and measurement considerations
Date Crossref
28/08/2026
Éditeur
Springer Science and Business Media LLC
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Simulation-Based Education in HealthcareArtificial Intelligence in Healthcare and EducationCardiac Arrest and Resuscitation

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.