When medical credentials conflict with stated accuracy: A factorial study of source credibility and answer revision in medical LLM interactions
Résumé fourni par la source
Abstract Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy. Author summary People increasingly ask conversational artificial intelligence systems to check a medical answer they have already formed. Such a second opinion is only useful if it stays independent after the system has heard what the person thinks. We wanted to know what happens when two things a user says about themselves disagree: their professional title and their recent track record. We put 480 Polish specialty examination questions to three widely used systems, ChatGPT, Claude and Gemini, and then challenged each answer in a separate conversation. The person challenging the answer was described either as a medical student or as an experienced specialist, and as having answered either two or eight of the last ten similar questions correctly. We found that the systems gave up a correct answer slightly more often for a specialist with a poor stated record than for a student with a good one, although the difference was small and not consistent across the three systems. The systems were not indiscriminately agreeable: they accepted correct corrections far more readily than incorrect ones. We take this to mean that agreement obtained after a user reveals a preferred answer should not be read as independent confirmation.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- When medical credentials conflict with stated accuracy: A factorial study of source credibility and answer revision in medical LLM interactions
- Date Crossref
- 01/09/2026
- Éditeur
- openRxiv
- Type
- posted-content
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.