Accès ouvert
2026
preprint
OpenAlex
Shuai Bai, Jiayong Deng, Sicheng Fan, Yikun Fu et autres
Computer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development through code and the command line. Real digital work requires both, interleaved rather than stacked end to end. We study hybrid CUAs that autonomously decide when to explore …
Accès ouvert
2026
preprint
OpenAlex
Dunjie Lu, Shuai Bai, Tianyi Bai, Sicheng Fan et autres
Native computer use offers a general interface for agents to operate almost any software available to people, but requires long-horizon state tracking, large-scale interactive experience, and learning from sparse yet verifiable outcomes. We introduce Qwen-CUA, a native computer-use agent with a 397B-A17B …
Accès ouvert
2026
preprint
OpenAlex
Dunjie Lu, Shuai Bai, Tianyi Bai, Sicheng Fan et autres
Native computer use offers a general interface for agents to operate almost any software available to people, but requires long-horizon state tracking, large-scale interactive experience, and learning from sparse yet verifiable outcomes. We introduce Qwen-CUA, a native computer-use agent with a 397B-A17B …
Accès ouvert
2026
preprint
OpenAlex
Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong, Weiming Wu et autres
Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents. We introduce OSWorld 2.0, a benchmark of 108 long-horizon computer-use workflows across everyday and professional tasks, …
Accès ouvert
2026
preprint
OpenAlex
Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong, Weiming Wu et autres
Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents. We introduce OSWorld 2.0, a benchmark of 108 long-horizon computer-use workflows across everyday and professional tasks, …
Accès ouvert
2026
preprint
OpenAlex
Bowen Wang, Dunjie Lu, Junli Wang, Tianyi Bai et autres
Reinforcement learning with verifiable rewards (RLVR) has driven breakthroughs in domains such as math, tool-use, and software engineering, yet its extension to computer-use agents (CUAs) has been bottlenecked by the scarcity of scalable training data with deterministic rewards. Constructing such data for …
Accès ouvert
2026
preprint
OpenAlex
Bowen Wang, Dunjie Lu, Junli Wang, Tianyi Bai et autres
Reinforcement learning with verifiable rewards (RLVR) has driven breakthroughs in domains such as math, tool-use, and software engineering, yet its extension to computer-use agents (CUAs) has been bottlenecked by the scarcity of scalable training data with deterministic rewards. Constructing such data for …
2025
conference-paper
OpenAlex
Yu Jiang, Dunjie Lu
cn
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Dunjie Lu, Yiheng Xu, Haoyuan Wu, Xinyuan Wang et autres
Training computer-use agents requires massive amounts of GUI interaction data, but manually annotating action trajectories at scale is prohibitively expensive. We present VideoAgentTrek, a scalable pipeline that automatically mines training data from publicly available screen-recorded videos at web scale, eliminating the need …
2025
conference-paper
OpenAlex
Xinyuan Wang, Bowen Wang, Dunjie Lu, Junlin Yang et autres
hk, cn, us, ca
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Yiheng Xu, Dunjie Lu, Zhi-Qiang Shen, Junli Wang et autres
Graphical User Interface (GUI) agents can automate complex tasks across digital environments, but their development is hindered by the scarcity of high-quality trajectory data for training. Existing approaches rely on expensive human annotation, making them unsustainable at scale. We propose AgentTrek, a …
Accès ouvert
2024
preprint
OpenAlex
Yiheng Xu, Zekun Wang, Junli Wang, Dunjie Lu et autres
Automating GUI tasks remains challenging due to reliance on textual representations, platform-specific action spaces, and limited reasoning capabilities. We introduce Aguvis, a unified vision-based framework for autonomous GUI agents that directly operates on screen images, standardizes cross-platform interactions and incorporates structured reasoning …