Accès ouvert
2026
preprint
OpenAlex
Haifeng Wang, Hua Wu, Tian Wu, Yu Sun et autres
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio. All modalities are trained from scratch under a unified next-group-of-tokens prediction objective, based on an ultra-sparse mixture-of-experts …
Accès ouvert
2026
preprint
OpenAlex
Haifeng Wang, Hua Wu, Tian Wu, Yu Sun et autres
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio. All modalities are trained from scratch under a unified next-group-of-tokens prediction objective, based on an ultra-sparse mixture-of-experts …