Accès ouvert
2026
preprint
OpenAlex
Jing Jiang, Yue Yang, Xinkai Jiang, Gedas Bertasius et autres
Robot manipulation policies are improving quickly, and real-robot evaluation remains the standard evidence for that progress. It still relies on a human to reset the scene between rollouts, which consumes operator time and leaves the initial state distribution unspecified, so results reproduce …
Accès ouvert
2026
preprint
OpenAlex
Yue Yang, Diego Romeres, Chiori Hori, Gedas Bertasius et autres
Vision-Language-Action (VLA) policies fuse multimodal sensory inputs, but training on limited and homogeneous robot demonstrations encourages spurious inter-sensor correlations rather than task-relevant signal, a failure we term modality entanglement. Under real-world occlusions and distractors, this manifests as nuisance sensitivity to corruption of …
Accès ouvert
2026
preprint
OpenAlex
Ce Zhang, Ziyang Wang, Yulu Pan, Oluwatumininu Oguntola et autres
Grounded long-video question answering (Grounded LVQA) requires answering a question about a long video while localizing the short evidence interval that supports the answer. Recent agentic methods frame this task as multi-turn exploration with a single crop_video(start, end) action, which supports coarse-to-fine …
us
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Baiqi Li, Ce Zhang, Y D Fang, Yue Yang et autres
A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layouts, object histories, and gestures that language leaves underspecified, yet today's manipulation benchmarks pair an instruction with a single …
Accès ouvert
2026
preprint
OpenAlex
Nurislam Tursynbek, Zhiqiang Lao, Heather Yu, Gedas Bertasius et autres
Recent text-to-video diffusion transformers generate visually compelling frames, yet still struggle with temporal coherence, often producing flickering, drifting, or unstable motion. We show that these failures leave a clear imprint inside the model: incoherent videos consistently exhibit irregular, fragmented temporal diagonals in …
Accès ouvert
2026
preprint
OpenAlex
Ziyang Wang, Yue Zhang, Shoubin Yu, Ce Zhang et autres
Next-generation visual assistants, such as smart glasses, embodied agents, and always-on life-logging systems, must reason over an entire day or more of continuous visual experience. In ultra-long videos, relevant information is sparsely distributed across hours or days, making memory a fundamental challenge: …
Accès ouvert
2026
preprint
OpenAlex
Yan-Bo Lin, Jonah Casebeer, Long Mai, Aniruddha Mahapatra et autres
Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal control. We introduce V2M-ZERO, a video-to-music generation approach that generates time-aligned music with disentangled time synchronization and semantic control (e.g., genre, mood) from video …
Accès ouvert
2026
conference-paper
OpenAlex
Xizi Wang, Feng Cheng, Ziyang Wang, Huiyu Wang et autres
Video temporal grounding aims to localize relevant temporal boundaries in a video given a textual prompt. Recent work has focused on enabling Video LLMs to perform video temporal grounding via next-token prediction of temporal timestamps. However, accurately localizing timestamps in videos remains …
us, mx
(code pays fourni par la source)
Accès ouvert
2026
conference-paper
OpenAlex
Yan-Bo Lin, Kevin Lin, Zhengyuan Yang, Linjie Li et autres
In this paper, we introduce zero-shot audio-video editing, a novel task that requires transforming original audio-visual content to align with a specified textual prompt without additional model training. To evaluate this task, we curate a benchmark dataset, AVED-Bench, designed explicitly for zero-shot …
us, gb
(code pays fourni par la source)
Accès ouvert
2026
conference-paper
OpenAlex
Ce Zhang, Yale Song, Ruta Desai, Michael L. Iuzzolino et autres
Visual Planning for Assistance (VPA) aims to predict a sequence of user actions required to achieve a specified goal based on a video showing the user’s progress. Although recent advances in multimodal large language models (MLLMs) have shown promising results in video …
us
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Yue Yang, Shuo Cheng, Yu Fang, Homanga Bharadhwaj et autres
General-purpose robots must master long-horizon manipulation, defined as tasks involving multiple kinematic structure changes (e.g., attaching or detaching objects) in unstructured environments. While Vision-Language-Action (VLA) models offer the potential to master diverse atomic skills, they struggle with the combinatorial complexity of sequencing …
Accès ouvert
2026
preprint
OpenAlex
Baiqi Li, Kangyi Zhao, Ce Zhang, Chancharik Mitra et autres
Fine-grained spatio-temporal understanding is essential for video reasoning and embodied AI. Yet, while Multimodal Large Language Models (MLLMs) master static semantics, their grasp of temporal dynamics remains brittle. We present TimeBlind, a diagnostic benchmark for compositional spatio-temporal understanding. Inspired by cognitive science, …