A Multi-AI Agent Framework for Interactive Neurosurgical Education and Evaluation: From Vignettes to Virtual Conversations
Rattachement africain : us. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
ABSTRACT Background and Objectives Traditional medical board examinations present clinical information in static vignettes with multiple-choices, fundamentally different from how physicians gather and integrate data in practice. Recent advances in Large Language Models (LLMs) offer promising approaches to creating more realistic clinical interactive conversations. However, these approaches are limited in neurosurgery, where patient communication capacity varies significantly and diagnosis heavily relies on objective data like imaging and neurological examinations. We aimed to develop and evaluate a multi-AI agent conversation framework for neurosurgical case assessment that enables realistic clinical interactions through simulated patients and structured access to objective clinical data. Methods We developed a framework to convert 608 Self-Assessment in Neurological Surgery (SANS) first-order diagnosis questions into conversation sessions using three specialized AI agents: Patient AI for subjective information, System AI for objective data, and Clinical AI for diagnostic reasoning. We evaluated GPT-4o’s diagnostic accuracy across traditional vignettes, patient-only conversations, and patient+system AI interactions, with human benchmark testing from ten neurosurgery residents. Results GPT-4o showed significant performance drops from traditional vignettes to conversational formats in both multiple-choice (89.0% to 60.9%, p<0.0001) and free-response scenarios (78.4% to 30.3%, p<0.0001). Adding access to objective data through System AI improved performance (to 67.4%, p=0.0015 and 61.8%, p<0.0001, respectively). Questions requiring image interpretation showed similar patterns but lower accuracy. Residents outperformed GPT-4o in free-response conversations (70.0% vs 28.3%, p=0.0030) using fewer interactions and reported high educational value of the interactive format. Conclusions This multi-AI agent framework provides both a more challenging evaluation method for LLMs and an engaging educational tool for neurosurgical training. The significant performance drops in conversational formats suggest that traditional multiple-choice testing may overestimate LLMs’ clinical reasoning capabilities, while the framework’s interactive nature offers promising applications for enhancing medical education.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- A Multi-AI Agent Framework for Interactive Neurosurgical Education and Evaluation: From Vignettes to Virtual Conversations
- Date Crossref
- 24/08/2025
- Éditeur
- openRxiv
- Type
- posted-content
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Hinge Health pays non établi dans la noticeÉtablissement de santé
-
NYU Langone Health pays non établi dans la noticeÉtablissement de santé
-
NYU Langone Department of Neurosurgery 550 First Ave New York pays non établi dans la noticeOrganisation à but non lucratif
Hinge Health, NYU Langone Health et NYU Langone Department of Neurosurgery 550 First Ave New York.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.