Aller au contenu principal
Accès ouvert déclaré 2025 preprint

A Multi-AI Agent Framework for Interactive Neurosurgical Education and Evaluation: From Vignettes to Virtual Conversations

0Citations signalées, ce qui n’est pas une note de qualité
2Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : us. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

ABSTRACT Background and Objectives Traditional medical board examinations present clinical information in static vignettes with multiple-choices, fundamentally different from how physicians gather and integrate data in practice. Recent advances in Large Language Models (LLMs) offer promising approaches to creating more realistic clinical interactive conversations. However, these approaches are limited in neurosurgery, where patient communication capacity varies significantly and diagnosis heavily relies on objective data like imaging and neurological examinations. We aimed to develop and evaluate a multi-AI agent conversation framework for neurosurgical case assessment that enables realistic clinical interactions through simulated patients and structured access to objective clinical data. Methods We developed a framework to convert 608 Self-Assessment in Neurological Surgery (SANS) first-order diagnosis questions into conversation sessions using three specialized AI agents: Patient AI for subjective information, System AI for objective data, and Clinical AI for diagnostic reasoning. We evaluated GPT-4o’s diagnostic accuracy across traditional vignettes, patient-only conversations, and patient+system AI interactions, with human benchmark testing from ten neurosurgery residents. Results GPT-4o showed significant performance drops from traditional vignettes to conversational formats in both multiple-choice (89.0% to 60.9%, p<0.0001) and free-response scenarios (78.4% to 30.3%, p<0.0001). Adding access to objective data through System AI improved performance (to 67.4%, p=0.0015 and 61.8%, p<0.0001, respectively). Questions requiring image interpretation showed similar patterns but lower accuracy. Residents outperformed GPT-4o in free-response conversations (70.0% vs 28.3%, p=0.0030) using fewer interactions and reported high educational value of the interactive format. Conclusions This multi-AI agent framework provides both a more challenging evaluation method for LLMs and an engaging educational tool for neurosurgical training. The significant performance drops in conversational formats suggest that traditional multiple-choice testing may overestimate LLMs’ clinical reasoning capabilities, while the framework’s interactive nature offers promising applications for enhancing medical education.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
A Multi-AI Agent Framework for Interactive Neurosurgical Education and Evaluation: From Vignettes to Virtual Conversations
Date Crossref
24/08/2025
Éditeur
openRxiv
Type
posted-content

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Hinge Health pays non établi dans la notice
    Établissement de santé
  • NYU Langone Health pays non établi dans la notice
    Établissement de santé
  • NYU Langone Department of Neurosurgery 550 First Ave New York pays non établi dans la notice
    Organisation à but non lucratif

Hinge Health, NYU Langone Health et NYU Langone Department of Neurosurgery 550 First Ave New York.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Artificial Intelligence in Healthcare and EducationSurgical Simulation and TrainingRadiology practices and education

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.