Benchmarking clinical reasoning and accuracy of large language models on breast oncology multiple-choice questions.
Roupen Odabashian, Ameer Basta, Ritu Sidgal, Trevor Lin et autres
e13637 Background: Large language models (LLMs) like GPT-4 (OpenAI) and Claude Opus (Anthropic) showed high accuracy in medical multiple-choice exams, but data on their oncology-specific clinical reasoning and performance is limited. This study evaluates their accuracy and clinical reasoning on breast oncology …
us, pk (code pays fourni par la source)