Adaptive Momentum Mixture-of-Experts for Continual Visual Question Answering
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Multimodal large language models (MLLMs) have attracted considerable attention for their impressive capabilities in understanding and generating visual-language content, particularly in tasks such as visual question answering (VQA). However, the rapid evolution of knowledge in real-world applications poses challenges for these models: offline training becomes increasingly costly, and exposure to non-stationary data streams often leads to catastrophic forgetting. In this paper, we propose CL-MoE+, a dual-momentum Mixture-of-Experts (MoE) framework based on MLLMs for continual VQA. Our method integrates continual learning into MLLMs to leverage the rich commonsense knowledge embedded in large language models.We introduce a Dual-Router MoE (RMoE) module that selects both global and local experts through task-level and instance-level routers, enabling robust and context-aware expert allocation. Furthermore, we design an adaptive Momentum MoE (MMoE) to update experts’ parameters based on the knowledge drift degree and their relevance to specific tasks, thereby facilitating knowledge integration without forgetting. Extensive experiments on a 10-task split of the VQA v2 benchmark demonstrate that CL-MoE+ achieves state-of-the-art performance, validating its effectiveness in both retaining historical knowledge and learning new information in the continual learning setting.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Adaptive Momentum Mixture-of-Experts for Continual Visual Question Answering
- Date Crossref
- 01/04/2026
- Éditeur
- Institute of Electrical and Electronics Engineers (IEEE)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
East China Normal University pays non établi dans la noticeUniversité ou école supérieure
-
Shanghai Open University Center of Open Distance Education pays non établi dans la noticeUniversité ou école supérieure
-
Fudan University pays non établi dans la noticeUniversité ou école supérieure
-
School of Computer Science and Technology pays non établi dans la noticeUniversité ou école supérieure
-
Zhejiang Zhuiqingting Technology Company Ltd. Zhejiang Zhuiqingting Technology Co. pays non établi dans la noticeEntreprise
-
College of Computer Science and Artificial Intelligence pays non établi dans la noticeUniversité ou école supérieure
East China Normal University, Center of Open Distance Education — Shanghai Open University et Fudan University, avec 3 autres affiliations.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.