Accès ouvert
2026
preprint
OpenAlex
Zijian Kan, Wei Wang, Long Luo, Bing Zhao et autres
Reward models play an essential role in aligning visual generative models, yet most existing visual reward models use a single scalar score or rely on fixed criteria that cannot adapt to different instructions. This limits both interpretability and task sensitivity, especially for …
Accès ouvert
2026
preprint
OpenAlex
Guangzheng Hu, Ziyue Jiang, Weixu Qiao, Lixin Zhang et autres
Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as "OmniJudges" for evaluation and automatic annotation. How reliably they understand what they score remains unclear, since existing benchmarks and training data tend to …
Accès ouvert
2026
preprint
OpenAlex
Shengyi Wang, Niantong Li, Guangzheng Hu, Hong Qi et autres
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary …
Accès ouvert
2026
preprint
OpenAlex
Shengyi Wang, Niantong Li, Guangzheng Hu, Hong Qi et autres
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary …
2026
article
OpenAlex
Congli Cui, Yuhang Bian, Weixu Qiao, Lijun Wang et autres
cn
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Xinyang He, Wei Wang, Bing Zhao, Xuan Ren et autres
Sparse autoencoders (SAEs) have emerged as a powerful tool for decomposing superposed language model representations into sparse and interpretable features. However, training SAEs is computationally expensive, and available open-source SAE models remain limited. In this work, we introduce \textbf{Qwen3-Instruct SAE}, a comprehensive …
Accès ouvert
2026
preprint
OpenAlex
Xinyang He, Wei Wang, Bing Zhao, Xuan Ren et autres
Sparse autoencoders (SAEs) have emerged as a powerful tool for decomposing superposed language model representations into sparse and interpretable features. However, training SAEs is computationally expensive, and available open-source SAE models remain limited. In this work, we introduce \textbf{Qwen3-Instruct SAE}, a comprehensive …
cn, us
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Niantong Li, Guangzheng Hu, Weixu Qiao, Ying Ba et autres
Text-to-Image generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple text-image alignment can no longer satisfy users' pressing demands for faithful real-world reconstruction and genuine creative expression. Existing benchmarks, however, remain anchored …
Accès ouvert
2026
preprint
OpenAlex
Niantong Li, Guangzheng Hu, Weixu Qiao, Ying Ba et autres
Text-to-Image generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple text-image alignment can no longer satisfy users' pressing demands for faithful real-world reconstruction and genuine creative expression. Existing benchmarks, however, remain anchored …
Accès ouvert
2026
preprint
OpenAlex
Bing Zhao, Chenfei Wu, Deqing Li, Hao Meng et autres
We present Qwen-Image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise image editing within a single framework. Despite recent progress, existing models still struggle with ultra-long text rendering, multilingual typography, high-resolution photorealism, robust instruction following, and efficient deployment, …
Accès ouvert
2026
preprint
OpenAlex
Bing Zhao, Chenfei Wu, Deqing Li, Hao Meng et autres
We present Qwen-Image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise image editing within a single framework. Despite recent progress, existing models still struggle with ultra-long text rendering, multilingual typography, high-resolution photorealism, robust instruction following, and efficient deployment, …
Accès ouvert
2026
preprint
OpenAlex
Xin An, Jingyi Cai, Xiangyang Chen, Huayao Liu et autres
Addressing the challenges of fragmented task definitions and the heterogeneity of unstructured data in multimodal parsing, this paper proposes the Omni Parsing framework. This framework establishes a Unified Taxonomy covering documents, images, and audio-visual streams, introducing a progressive parsing paradigm that bridges …