Accès ouvert
2026
preprint
OpenAlex
Jiacheng Ruan, Daize Dong, Xiaoye Qu, Tong Zhu et autres
Mixture-of-Experts (MoE) models substantially improve performance by increasing the capacity of dense architectures. However, directly training MoE models requires considerable computational resources and introduces extra overhead in parameter storage and deployment. Therefore, it is critical to develop an approach that leverages the …
Accès ouvert
2026
preprint
OpenAlex
Zefeng He, Siyuan Huang, Xiaoye Qu, Yafu Li et autres
Recent multimodal generation models have achieved remarkable progress on general-purpose generation tasks, yet continue to struggle with complex instructions and specialized downstream tasks. Inspired by the success of advanced agent frameworks such as Claude Code, we propose \textbf{GEMS} (Agent-Native Multimodal \textbf{GE}neration with …
Accès ouvert
2026
preprint
OpenAlex
Jiacheng Ruan, Daize Dong, Xiaoye Qu, Tong Zhu et autres
Mixture-of-Experts (MoE) models substantially improve performance by increasing the capacity of dense architectures. However, directly training MoE models requires considerable computational resources and introduces extra overhead in parameter storage and deployment. Therefore, it is critical to develop an approach that leverages the …
cn
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Yicheng Zou, Dongsheng Zhu, Lin Zhu, Tong Zhu et autres
We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancement across both general and scientific domains. Beyond stronger reasoning and image-text understanding capabilities, its intelligence is augmented with advanced agent capabilities. …
Accès ouvert
2026
preprint
OpenAlex
Yicheng Zou, Dongsheng Zhu, Lin Zhu, Tong Zhu et autres
We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancement across both general and scientific domains. Beyond stronger reasoning and image-text understanding capabilities, its intelligence is augmented with advanced agent capabilities. …
Accès ouvert
2026
conference-paper
OpenAlex
Yonghong Jia, Yuntao Du, Kailin Jiang, Y. T. Liang et autres
Large Multimodal Models (LMMs) face notable challenges when encountering multimodal knowledge conflicts, particularly under retrieval-augmented generation (RAG) frameworks, where the contextual information from external sources may contradict the model’s internal parametric knowledge, leading to unreliable outputs. However, existing benchmarks fail to reflect …
cn, hk
(code pays fourni par la source)
Accès ouvert
2026
conference-paper
OpenAlex
Xiang Fang, Wanlong Fang, Changshuo Wang, Xiaoye Qu et autres
Driven by the wave of large language models, Video-Language Models (VLMs) have become a significant yet challenging technology to bridge the gap between videos and texts. Although previous VLM works have made significant progress, almost all of them implicitly assume that all …
Accès ouvert
2026
preprint
OpenAlex
Guanyu Jiang, Zhaochen Su, Xiaoye Qu, Yi R
Multimodal agents can now tackle complex reasoning tasks with diverse tools, yet they still suffer from inefficient tool use and inflexible orchestration in open-ended settings. A central challenge is enabling such agents to continually improve without parameter updates by learning from past …