Accès ouvert
2026
preprint
OpenAlex
Yicheng Zou, Dongsheng Zhu, Lin Zhu, Tong Zhu et autres
We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancement across both general and scientific domains. Beyond stronger reasoning and image-text understanding capabilities, its intelligence is augmented with advanced agent capabilities. …
Accès ouvert
2026
preprint
OpenAlex
Yicheng Zou, Dongsheng Zhu, Lin Zhu, Tong Zhu et autres
We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancement across both general and scientific domains. Beyond stronger reasoning and image-text understanding capabilities, its intelligence is augmented with advanced agent capabilities. …
Accès ouvert
2026
preprint
OpenAlex
Yechen Zhang, Shuhao Xing, Junhao Huang, Kai Lv et autres
Recent advances in spectral optimization, notably Muon, have demonstrated that constraining update steps to the Stiefel manifold can significantly accelerate training and improve generalization. However, Muon implicitly assumes an isotropic optimization landscape, enforcing a uniform spectral update norm across all eigen-directions. We …
Accès ouvert
2026
preprint
OpenAlex
Yechen Zhang, Shuhao Xing, Junhao Huang, Kai Lv et autres
Recent advances in spectral optimization, notably Muon, have demonstrated that constraining update steps to the Stiefel manifold can significantly accelerate training and improve generalization. However, Muon implicitly assumes an isotropic optimization landscape, enforcing a uniform spectral update norm across all eigen-directions. We …
cn
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Yunhua Zhou, Junhao Huang, Shuhao Xin, Yechen Zhang et autres
The concept of Critical Batch Size, as pioneered by OpenAI, has long served as a foundational principle for large-scale pre-training. However, with the paradigm shift towards the Warmup-Stable-Decay (WSD) learning rate scheduler, we observe that the original theoretical framework and its underlying …
Accès ouvert
2026
preprint
OpenAlex
Yunhua Zhou, Junhao Huang, Shuhao Xin, Yechen Zhang et autres
The concept of Critical Batch Size, as pioneered by OpenAI, has long served as a foundational principle for large-scale pre-training. However, with the paradigm shift towards the Warmup-Stable-Decay (WSD) learning rate scheduler, we observe that the original theoretical framework and its underlying …
cn
(code pays fourni par la source)