Benefits of Chain-of-Thought Prompting for Clinical Record Rubric Evaluation in Undergraduate Medical Education: Experimental Evaluation Study With Medical Faculty
Rattachement africain : es. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Background: Large language models in artificial intelligence have been among the tools with a significant and real impact on people's daily lives. In this regard, they serve as an aid in specific fields, such as education, helping educators with cumbersome tasks such as periodic evaluations. Objective: This study focused on analyzing the benefits of large language models, particularly the chain-of-thought (CoT) strategy, for the task of evaluating students' Spanish-language medical record writing. The aim was 2-fold: first, we attempted to save time and resources, and second, we used the reasoning of the CoT strategy to evaluate the rubrics and their interpretations. Methods: The proposed solution assessed the application of 2 models-Llama 3.1 and Claude 3.5-in combination with one-shot and CoT to evaluate how medical students write medical records in Spanish. First, machine learning metrics were applied to measure the performance of the solutions. Then, different statistical analyses were performed at the clinical record and item levels. Finally, differences between the proposed models and evaluators were studied in depth. Results: A maximum of 3807 items were evaluated. Claude obtained the best accuracy with slight differences between one-shot and CoT (86.4% and 85.0%, respectively). However, Claude with CoT outperformed the rest of the combinations on all complementary metrics, initially achieving a sensitivity of 94.2%, specificity of 59.5%, precision of 85.8%, and F1-score of 89.6%. Expert review of CoT reasoning determined that 63.8% of the discrepancies were model hits, raising Claude's final accuracy to 94.6% (SD 4.3%). In the final phase, sensitivity was 98.0% (SD 2.3%), specificity improved to 83.3% (SD 14.3%), and F1-score reached 96.2% (SD 3.3%). Sectional analysis showed greater difficulties in the "History of present illness" section (n=125 discordances). Conclusions: CoT demonstrated strong potential for supporting the evaluation of clinical records written in Spanish by medical students and providing feedback to them. More importantly, it showed significant promise in assisting professors by assessing the quality of their rubrics and identifying possible errors.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Benefits of Chain-of-Thought Prompting for Clinical Record Rubric Evaluation in Undergraduate Medical Education: Experimental Evaluation Study With Medical Faculty
- Date Crossref
- 23/07/2026
- Éditeur
- JMIR Publications Inc.
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.