Automated Generation and Human Evaluation of Neurosurgical Board Examination Self-Assessment Questions
Rattachement africain : us. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
BACKGROUND AND OBJECTIVES: Multiple-choice questions are the primary assessment format for neurosurgical board certification. Creating high-quality examination questions requires significant expert time and resources. The goal of this study was to develop an automated system to generate board-style neurosurgical multiple-choice questions using state-of-the-art vision-language models and compare their quality with authentic self-assessment questions. METHODS: We developed an automated pipeline using OpenAI generative pre-trained transformer (GPT)-4o and Anthropic Claude Sonnet-3.5 to generate neurosurgical board-style questions from Neurosurgery Publications articles. We generated 89 587 synthetic questions: 45 689 with GPT-4o and 43 898 with Claude. Each question was associated with a single image extracted from the articles' figures. We evaluated the quality of synthetic questions through 5 surveys comparing 20 synthetic questions (10 from each model) with 10 authentic questions from the Self-Assessment for Neurological Surgeons (SANS) question bank. Each survey was completed by a neurosurgery resident and an attending who guessed the source [human vs artificial intelligence (AI)-generated] and rated suitability for board examination use. We also evaluated the question-answering performance of the generalist GPT-4o and the specialized CNS-Obsidian. RESULTS: SANS questions were more often perceived as human-made than GPT-generated (residents, P = .0002; attendings, P = .1091) and Claude-generated (residents, P = .0002; attendings, P = .0272) questions. Notably, 54% of AI-generated questions misled at least one evaluator, and 23% misled both. In quality assessments, SANS questions outperformed GPT-generated (residents, P < 10 −5 ; attendings, P = .0001) and Claude-generated (residents and attendings, P < 10 −5 ) questions. Particularly, 25% of AI-generated questions were rated as suitable for board examinations vs 72% of human-generated questions when measured by evaluator consensus ( P < 10 −7 ). CONCLUSION: Although quality gaps exist between AI-generated and human-created neurosurgical board examination questions, our approach demonstrates the potential of vision-language models to augment assessment development in specialized medical fields, reducing the burden on examination boards and credentialing organizations.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Automated Generation and Human Evaluation of Neurosurgical Board Examination Self-Assessment Questions
- Date Crossref
- 23/07/2026
- Éditeur
- Ovid Technologies (Wolters Kluwer Health)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Washington University in St. Louis pays non établi dans la noticeUniversité ou école supérieure
-
Neurological Surgery pays non établi dans la noticeÉtablissement de santé
-
Barrow Neurological Institute Department of Neurosurgery pays non établi dans la noticeÉtablissement de santé
-
NYU Langone Health Department of Neurological Surgery pays non établi dans la noticeOrganisation à but non lucratif
-
Washington University in Saint Louis Department of Neurosurgery pays non établi dans la noticeUniversité ou école supérieure
Washington University in St. Louis, Neurological Surgery et Department of Neurosurgery — Barrow Neurological Institute, avec 2 autres affiliations.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.