Accès ouvert
2026
preprint
OpenAlex
Fay Elhassan, David Sasu, Alexandra Kulinkina, Lars Klein et autres
We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform collecting blinded pairwise preferences alongside multi-criterion rubric ratings. Clinicians …
Accès ouvert
2026
preprint
OpenAlex
Fay Elhassan, David Sasu, Alexandra Kulinkina, Lars Klein et autres
We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform collecting blinded pairwise preferences alongside multi-criterion rubric ratings. Clinicians …
ch
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Yusuf Kesmen, Fay Elhassan, Jiayi Ma, Julien Stalhandske et autres
Large language models (LLMs) are increasingly used for conversational clinical decision support, yet they conflate next token prediction with probabilistic decision making. We argue that this conflation reflects an architectural limitation: such systems lack explicit posterior tracking, controllable abstention thresholds, and auditable …
Accès ouvert
2026
preprint
OpenAlex
Yusuf Kesmen, Fay Elhassan, Jiayi Ma, Julien Stalhandske et autres
Large language models (LLMs) are increasingly used for conversational clinical decision support, yet they conflate next token prediction with probabilistic decision making. We argue that this conflation reflects an architectural limitation: such systems lack explicit posterior tracking, controllable abstention thresholds, and auditable …
ch, dk
(code pays fourni par la source)