Aller au contenu principal
Accès ouvert déclaré 2026 article

Benefits of Chain-of-Thought Prompting for Clinical Record Rubric Evaluation in Undergraduate Medical Education: Experimental Evaluation Study With Medical Faculty

0Citations signalées, ce qui n’est pas une note de qualité
2Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : es. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Background: Large language models in artificial intelligence have been among the tools with a significant and real impact on people's daily lives. In this regard, they serve as an aid in specific fields, such as education, helping educators with cumbersome tasks such as periodic evaluations. Objective: This study focused on analyzing the benefits of large language models, particularly the chain-of-thought (CoT) strategy, for the task of evaluating students' Spanish-language medical record writing. The aim was 2-fold: first, we attempted to save time and resources, and second, we used the reasoning of the CoT strategy to evaluate the rubrics and their interpretations. Methods: The proposed solution assessed the application of 2 models-Llama 3.1 and Claude 3.5-in combination with one-shot and CoT to evaluate how medical students write medical records in Spanish. First, machine learning metrics were applied to measure the performance of the solutions. Then, different statistical analyses were performed at the clinical record and item levels. Finally, differences between the proposed models and evaluators were studied in depth. Results: A maximum of 3807 items were evaluated. Claude obtained the best accuracy with slight differences between one-shot and CoT (86.4% and 85.0%, respectively). However, Claude with CoT outperformed the rest of the combinations on all complementary metrics, initially achieving a sensitivity of 94.2%, specificity of 59.5%, precision of 85.8%, and F1-score of 89.6%. Expert review of CoT reasoning determined that 63.8% of the discrepancies were model hits, raising Claude's final accuracy to 94.6% (SD 4.3%). In the final phase, sensitivity was 98.0% (SD 2.3%), specificity improved to 83.3% (SD 14.3%), and F1-score reached 96.2% (SD 3.3%). Sectional analysis showed greater difficulties in the "History of present illness" section (n=125 discordances). Conclusions: CoT demonstrated strong potential for supporting the evaluation of clinical records written in Spanish by medical students and providing feedback to them. More importantly, it showed significant promise in assisting professors by assessing the quality of their rubrics and identifying possible errors.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Benefits of Chain-of-Thought Prompting for Clinical Record Rubric Evaluation in Undergraduate Medical Education: Experimental Evaluation Study With Medical Faculty
Date Crossref
23/07/2026
Éditeur
JMIR Publications Inc.
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Clinical Reasoning and Diagnostic SkillsInnovations in Medical EducationEducational Assessment and Improvement

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.