Aller au contenu principal
2025 conference-abstract

S2998 Evaluating the Accuracy of ChatGPT on MKSAP Gastroenterology and Hepatology Board Review Questions

0Citations signalées, ce qui n’est pas une note de qualité
0Institutions déclarées
0Pays d’affiliation déclarés

Le résumé fourni par la source

Introduction: Large language models including ChatGPT can augment medical education and clinical decision-making. There are various models of ChatGPT, with some having better capabilities for advanced reasoning. To date, ChatGPT’s ability to answer board-style questions in Gastroenterology (GI) and Hepatology categories has not been evaluated. This study aims to assess the accuracy of 2 different models of ChatGPT, 4o and o3, in answering board review questions from the American College of Physicians’ Medical Knowledge Self-Assessment Program (MKSAP), focusing specifically on GI and hepatology. Methods: All 155 questions in the GI and Hepatology section of ACP MKSAP were selected for this study. Each question with associated imaging and lab data were copied into ChatGPT 4o, the model designed for most tasks, and then ChatGPT o3, OpenAI’s reasoning model designed for complex tasks requiring deep reasoning. The options selected by each ChatGPT model were recorded and compared with the answers provided by MKSAP. Questions were divided into the following categories: general gastroenterology, inflammatory bowel disease (IBD), hepatobiliary, GI cancer, and esophageal. Chi square analyses were performed. Results: ChatGPT o3, the advanced deep reasoning model, significantly outperformed ChatGPT 4o, with 96% of questions answered correctly compared to 85% (P = 0.009). By category, ChatGPT o3 answered 100% of IBD, 96% of hepatobiliary, 91% GI cancer, 86% esophageal, and 97% general GI questions correctly. ChatGPT 4o answered 100% IBD, 89% hepatobiliary, 82% GI cancer, 79% esophageal, and 81% general GI questions correctly. There was a significant difference between the 2 models for general GI questions (P = 0.005). Conclusion: Overall, ChatGPT had a high but not perfect accuracy rate in answering GI and hepatology internal medicine board review questions. ChatGPT o3 significantly outperformed ChatGPT 4o. This may be attributed to the o3 model’s reinforcement learning approach which enhances its capabilities in open-ended situations, particularly those involving visuals and multi-step workflows. More specifically, the o3 model had higher accuracy for general GI questions, but not for other GI subcategories, compared to the 4o model. Future directions for this study include comparing ChatGPT’s accuracy to that of internal medicine residents along with internal medicine attendings and GI attendings.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
S2998 Evaluating the Accuracy of ChatGPT on MKSAP Gastroenterology and Hepatology Board Review Questions
Date Crossref
01/10/2025
Éditeur
Ovid Technologies (Wolters Kluwer Health)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les sujets associés

Artificial Intelligence in Healthcare and EducationClinical Reasoning and Diagnostic SkillsRadiomics and Machine Learning in Medical Imaging

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.