Aller au contenu principal
Accès ouvert déclaré 2026 article

Cross-modal bias in medical vision-language models: a pipeline-aware framework for mechanisms, evaluation, and mitigation

0Citations signalées, ce qui n’est pas une note de qualité
4Institutions déclarées
3Pays d’affiliation déclarés

Rattachement africain : bd, us, cn. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Medical vision-language models encode images and clinical text in a shared representation. Across radiology and ophthalmology, their diagnostic performance now approaches that of specialist clinicians. The mechanism behind that performance is also the source of a problem that has gone largely unexamined. These models are trained by contrastive alignment, so bias from the image encoder and bias from the text encoder meet at a single point: the alignment interface. There they can interact and compound in ways that single-modality systems never experience. Fairness research has so far studied the two modalities separately, and medical vision-language models have fallen into the gap between those literatures. We organize the evidence into a three-tier taxonomy keyed to the pretraining pipeline. Tier 1 is data-level bias in the pretraining corpus. Tier 2 is alignment bias produced at the contrastive interface. Tier 3 is inference-time bias that surfaces during deployment. Within this structure, we compare the major evaluation benchmarks, show where they disagree, and assign each mitigation strategy to the tier it actually addresses. Three points emerge. First, no published method spans all three tiers; mitigation is fragmented by pipeline stage. Second, fine-tuning does not remove alignment-stage bias, even in parameter-efficient form, which shifts the burden of debiasing onto pretraining rather than adaptation. Third, the inference-time failures are more dangerous than the literature suggests. Medical-specialist models will abandon a correct reading and defer to a confident user on most trials, and specialization appears to make this worse, not better. We close with a research agenda.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Cross-modal bias in medical vision-language models: a pipeline-aware framework for mechanisms, evaluation, and mitigation
Date Crossref
07/08/2026
Éditeur
Frontiers Media SA
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Bangladesh University of Professionals Department of Information and Communication Technology pays non établi dans la notice
    Université ou école supérieure
  • California Lutheran University pays non établi dans la notice
    Université ou école supérieure
  • ElSohly Laboratories (United States) pays non établi dans la notice
    Entreprise
  • Cloud Computing Center pays non établi dans la notice
    Structure de recherche
  • ELITE Research Lab pays non établi dans la notice
    Structure de recherche
  • Faculty of Information Science & Technology Centre of Excellence (COE) of Advanced Cloud pays non établi dans la notice
    Université ou école supérieure

Department of Information and Communication Technology — Bangladesh University of Professionals, California Lutheran University et ElSohly Laboratories (United States), avec 3 autres affiliations.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Artificial Intelligence in Healthcare and EducationMultimodal Machine Learning ApplicationsDomain Adaptation and Few-Shot Learning

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.