Qwen-Image-2.0-RL Technical Report
Yan Xu, Kaiyuan Gao, Yuxiang Chen, Yilei Chen et autres
We present Qwen-Image-2.0-RL, a post-training pipeline that applies reinforcement learning from human feedback (RLHF) and on-policy distillation (OPD) to improve both the visual quality and instruction-following capability of the Qwen-Image-2.0 diffusion model. To provide reliable reward signals, we construct task-specific composite reward …