Accès ouvert
2026
preprint
OpenAlex
Zitai Huang, Taiyi Su, Jian Zhu, Jianjun Zhang et autres
Long-horizon robot manipulation requires not only stable local visuomotor control, but also continuous target tracking and reliable task progress assessment throughout execution. This challenge becomes particularly critical when multiple objects share identical appearances and must be manipulated in a prescribed order. In …
Accès ouvert
2026
preprint
OpenAlex
Tianbin Liu, Jian Zhu, Taiyi Su, Jianjun Zhang et autres
FastWAM-style world action models enable efficient action-only inference, but generalize poorly under visual distribution shifts. Their reconstruction-oriented representations emphasize appearance-specific details, limiting generalization to unseen scenes and objects. Without observation history, the model also lacks temporal evidence for robustly identifying task-relevant state …
Accès ouvert
2026
preprint
OpenAlex
Yi Xu, Yifan Hou, Xiaoyu Zhang
Video lectures are valuable educational resources, but their dense and lengthy formats often overwhelm novice learners. This difficulty stems from a fundamental pedagogical mismatch: while videos deliver transient information linearly, human learning requires constructing interconnected cognitive networks, a task that induces severe …
Accès ouvert
2026
preprint
OpenAlex
Yi Xu, Ruichao Hou, Tongwei Ren, Gangshan Wu
Salient object detection in optical remote sensing images (ORSI-SOD) requires dense predictions that preserve object completeness and structural continuity under complex backgrounds, scale variation, and irregular object shapes. Existing methods often localize salient regions, but their predictions may still suffer from structural …
Accès ouvert
2026
preprint
OpenAlex
Linhan Wang, Zijian An, Mingyuan Zhang, Chen Dai et autres
Video generative models provide rich physical priors for robot learning, yet existing world-action models (WAMs) face a fundamental trade-off: synchronous video generation at control rate is latency-prohibitive, while abandoning test-time visual imagination sacrifices task success. We show that visual imagination achieves both …
Accès ouvert
2026
preprint
OpenAlex
Tianbin Liu, Jian Zhu, Taiyi Su, Jianjun Zhang et autres
World Action Models (WAMs) have demonstrated strong robotic manipulation capabilities by augmenting pretrained video generative models with action experts. However, current WAMs still show limited instruction-following ability when conditioned solely on text instructions. We argue that this limitation stems in part from …
Accès ouvert
2026
preprint
OpenAlex
J ZHU, Jianjun Zhang, Taiyi Su, Tianbin Liu et autres
World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learning. Existing WAMs excel at physically grounded execution, but typically lack the explicit language-level planning interface in VLM-based VLAs …
Accès ouvert
2026
preprint
OpenAlex
Jianjun Zhang, J ZHU, Taiyi Su, C W et autres
World Action Models (WAMs) have shown strong potential for robotic manipulation by jointly modeling visual future dynamics and executable action sequences. However, existing video-action co-training methods primarily optimize appearance-oriented video latents, which may insufficiently capture the temporally evolving geometry required for precise …