Accès ouvert
2026
preprint
OpenAlex
Rong Jin, Jianming Ma, Yue Gao
Humanoid robots must traverse cluttered obstacle fields using onboard proprioceptive and visual observations, yet existing methods usually process multimodal observations without explicitly considering their different characteristics: proprioceptive observations are low-dimensional but governed by highly nonlinear robot dynamics, while egocentric visual observations are …
Accès ouvert
2026
preprint
OpenAlex
Jianming Ma, Rong Jin, Xiaxi Si, Yang Zhang et autres
Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in general-purpose robotic manipulation, yet their generated actions may violate hard physical constraints and therefore be unsafe or infeasible for deployment. Existing safety approaches either optimize statistical safety objectives without deterministic per-step …
Accès ouvert
2026
preprint
OpenAlex
AgiBot Research Team, Renhang Liu, Wenzhi Zhao, Zhuo Yang et autres
World-action models (WAM) predict future states to guide robot actions, enabling learning from both action-free video and action-labeled interaction. Most inherit pretrained video generators, leaving WAM pretraining and scaling underexplored. We introduce Genie Envisioner Act 2.0 (GE-Act 2.0), a world-action model whose …
Accès ouvert
2026
conference-paper
OpenAlex
Pengfei Zhou, Liliang Chen, Shengcong Chen, Di Chen et autres
Specifying robotic manipulation tasks in a manner that is both expressive and precise remains a central challenge.While visual goals provide a compact and unambiguous task specification, existing goal-conditioned policies often struggle with long-horizon manipulation due to their reliance on single-step action prediction …
Accès ouvert
2026
preprint
OpenAlex
Pengfei Zhou, Shengcong Chen, Di Chen, Jiaxu Wang et autres
Robotic manipulation requires models that generate executable actions while anticipating and evaluating their future consequences before physical execution. We present $τ_0$-World Model ($τ_0$-WM), a unified video-action world model that integrates policy learning, video prediction, and action evaluation within a single future-predictive framework. …
Accès ouvert
2026
preprint
OpenAlex
Pengfei Zhou, Shengcong Chen, Di Chen, Jiaxu Wang et autres
Robotic manipulation requires models that generate executable actions while anticipating and evaluating their future consequences before physical execution. We present $τ_0$-World Model ($τ_0$-WM), a unified video-action world model that integrates policy learning, video prediction, and action evaluation within a single future-predictive framework. …
Accès ouvert
2026
conference-paper
OpenAlex
Buqing Nie, Yang Zhang, Rong Jin, Zhanxiang Cao et autres
The human nervous system exhibits bilateral symmetry, enabling coordinated and balanced movements. However, existing Deep Reinforcement Learning (DRL) methods for humanoid robots neglect morphological symmetry of the robot, leading to uncoordinated and suboptimal behaviors. Inspired by human motor control, we propose Symmetry …
cn
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Pengfei Zhou, Liliang Chen, Shengcong Chen, Di Chen et autres
Specifying robotic manipulation tasks in a manner that is both expressive and precise remains a central challenge. While visual goals provide a compact and unambiguous task specification, existing goal-conditioned policies often struggle with long-horizon manipulation due to their reliance on single-step action …
Accès ouvert
2025
preprint
OpenAlex
DeepSeek-AI, Aixin Liu, Aoxue Mei, Bing Xue et autres
We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. The key technical breakthroughs of DeepSeek-V3.2 are as follows: (1) DeepSeek Sparse Attention (DSA): We introduce DSA, an efficient attention mechanism that substantially reduces computational complexity …
Accès ouvert
2025
preprint
OpenAlex
Hyunin Lee, Yong Zhang, Hoang Vu Nguyen, Xiaoyi Liu et autres
Cross-domain sequential recommendation (CDSR) aims to align heterogeneous user behavior sequences collected from different domains. While cross-attention is widely used to enhance alignment and improve recommendation performance, its underlying mechanism is not fully understood. Most researchers interpret cross-attention as residual alignment, where …
Accès ouvert
2025
article
OpenAlex
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song et autres
Abstract General reasoning represents a long-standing and formidable challenge in artificial intelligence (AI). Recent breakthroughs, exemplified by large language models (LLMs) 1,2 and chain-of-thought (CoT) prompting 3 , have achieved considerable success on foundational reasoning tasks. However, this success is heavily contingent …
cn
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Buqing Nie, Rong Jin, Zhanxiang Cao, Huangxuan Lin et autres
The human nervous system exhibits bilateral symmetry, enabling coordinated and balanced movements. However, existing Deep Reinforcement Learning (DRL) methods for humanoid robots neglect morphological symmetry of the robot, leading to uncoordinated and suboptimal behaviors. Inspired by human motor control, we propose Symmetry …