Aller au contenu principal
Accès ouvert déclaré 2026 preprint

Adaptive Multimodal Fusion in Radiology: Dynamic Balancing of Visual Findings and Clinical Context

0Citations signalées — pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Résumé fourni par la source

Abstract The automatic detection of thoracic pathologies remains a challenge when the visual evidence present in the chest X-ray (CXR) is subtle, ambiguous, or practically imperceptible, and the available textual information is limited, especially in the early phases of patient care. To address this limitation, in this work a clinically motivated subset of MIMIC-CXR is taken as the study set, composed of four diseases associated with dyspnea presentations in the emergency department (heart failure, pulmonary infection, COPD/asthma, and pulmonary embolism), along with a control class. This selection allows evaluating the multimodal fusion of visual and textual information in a clinically diverse set where the contribution of each modality can vary depending on the considered pathology and the quality of the available information. In this work, a comparative study of different combinations of visual and textual encoders is presented, with the objective of identifying the most suitable configuration for multimodal fusion. Based on this evaluation, we propose Neural Gated Fusion , a GMU-based architecture that incorporates an adapted Gated Multimodal Unit to regulate the contribution of visual and textual representations, in contrast to strategies based on static feature concatenation or late decision-level fusion. This architecture uses a neural gate to dynamically regulate the contribution of the visual representation of the radiograph and the textual representation of the clinical report before the multi-label classification stage. The experiments conducted on 25,245 multimodal pairs show that the ViT + ClinicalBERT combination with modulated fusion achieves a macro ROC-AUC of 0.8545 and a macro F1-score of 0.6528, outperforming unimodal models, decision-level fusion, and static early fusion. The per-class analysis suggests that the designed fusion strategy especially improves performance in pathologies with limited visual evidence, such as pulmonary embolism, without compromising the detection of normal cases. These results suggest that the dynamic regulation between image and clinical context is an effective strategy for improving the robustness of multimodal thoracic pathology detection systems. Author summary Computer-aided diagnosis systems typically rely only on visual data, like chest X-rays, to detect diseases. However, for conditions with subtle or overlapping visual signs (e.g., pulmonary embolism or heart failure), radiologists also rely heavily on patient clinical notes. Our study bridges this gap by creating an AI framework that mimics this real-world diagnostic process. We evaluated multiple visual and text analysis models to find the optimal combination. We then introduced ”Neural Gated Fusion,” an architecture that acts as a smart filter, dynamically deciding whether to prioritize the X-ray image or the clinical report based on the specific case. Tested on over 25,000 multimodal clinical records, our system significantly outperformed traditional single-data models, particularly in diagnosing hard-to-see conditions, without sacrificing accuracy on normal cases. This dynamic approach offers a more transparent and robust tool for clinical decision-making, moving artificial intelligence closer to the holistic reasoning used by medical professionals.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Adaptive Multimodal Fusion in Radiology: Dynamic Balancing of Visual Findings and Clinical Context
Date Crossref
07/09/2026
Éditeur
openRxiv
Type
posted-content

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

COVID-19 diagnosis using AIDomain Adaptation and Few-Shot LearningMultimodal Machine Learning Applications

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.