Accès ouvert
2026
preprint
OpenAlex
Songchun Zhang, Yaowei Li, Junhao Zhuang, Weiyang Jin et autres
We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person …
Accès ouvert
2026
preprint
OpenAlex
Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li et autres
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant …
Accès ouvert
2026
article
OpenAlex
Jie Huang, Yijun Liu, Haojun Xu, Yuming Li et autres
Accès ouvert
2026
preprint
OpenAlex
Yijun Liu, J W Huang, Zeyue Xue, Yuming Li et autres
Reward models guide text-to-image (T2I) systems toward outputs aligned with human preferences. However, typical reward models such as HPSv3 are trained on pre-annotated data from earlier T2I models, without accounting for quality discriminative shifts arising from evolving model capabilities and reinforcement learning …
Accès ouvert
2026
preprint
OpenAlex
Luxury, Jie Huang, Zihao Fan, Xiaoxiao Ma et autres
While recent autoregressive video diffusion models achieve remarkable streaming quality, they remain confined to low resolutions (e.g., 480P), leaving efficient, scalable, real-time high-resolution video generation a fundamental open challenge. To bridge this gap, we present Ultra Flash, a cascaded streaming framework capable …
Accès ouvert
2026
preprint
OpenAlex
Kan Huang, Yijun Liu, Jihao Liu, Zhiliang Lin