Cross-modal bias in medical vision-language models: a pipeline-aware framework for mechanisms, evaluation, and mitigation
Rafid Mehda, Ramisa Anjum Oishi, Tamzid Tanvi Alam, Md Kishor Morol et autres
Medical vision-language models encode images and clinical text in a shared representation. Across radiology and ophthalmology, their diagnostic performance now approaches that of specialist clinicians. The mechanism behind that performance is also the source of a problem that has gone largely unexamined. …
bd, us, cn (code pays fourni par la source)