Aller au contenu principal
Profil bibliographique

Gedas Bertasius

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

125Publications signalées
4419Citations signalées
2Affiliations récentes

Les institutions déclarées

Les domaines associés

Multimodal Machine Learning ApplicationsHuman Pose and Action RecognitionVideo Analysis and SummarizationDomain Adaptation and Few-Shot LearningAdvanced Image and Video Retrieval Techniques

Les publications récentes

Accès ouvert 2026 preprint OpenAlex

From Rollout to Reset: A Graph-Based Harness for Autonomous Long-Horizon Manipulation Evaluation

Jing Jiang, Yue Yang, Xinkai Jiang, Gedas Bertasius et autres

Robot manipulation policies are improving quickly, and real-robot evaluation remains the standard evidence for that progress. It still relies on a human to reset the scene between rollouts, which consumes operator time and leaves the initial state distribution unspecified, so results reproduce …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Sensing Which Modality Matters: Evidence-Gated Regularization for Robust VLA Policies

Yue Yang, Diego Romeres, Chiori Hori, Gedas Bertasius et autres

Vision-Language-Action (VLA) policies fuse multimodal sensory inputs, but training on limited and homogeneous robot demonstrations encourages spurious inter-sensor correlations rather than task-relevant signal, a failure we term modality entanglement. Under real-world occlusions and distractors, this manifests as nuisance sensitivity to corruption of …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA

Ce Zhang, Ziyang Wang, Yulu Pan, Oluwatumininu Oguntola et autres

Grounded long-video question answering (Grounded LVQA) requires answering a question about a long video while localizing the short evidence interval that supports the answer. Recent agentic methods frame this task as multi-turn exploration with a single crop_video(start, end) action, which supports coarse-to-fine …

us (code pays fourni par la source)

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation

Baiqi Li, Ce Zhang, Y D Fang, Yue Yang et autres

A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layouts, object histories, and gestures that language leaves underspecified, yet today's manipulation benchmarks pair an instruction with a single …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

TeDiO: Temporal Diagonal Optimization for Training-Free Coherent Video Diffusion

Nurislam Tursynbek, Zhiqiang Lao, Heather Yu, Gedas Bertasius et autres

Recent text-to-video diffusion transformers generate visually compelling frames, yet still struggle with temporal coherence, often producing flickering, drifting, or unstable motion. We show that these failures leave a clear imprint inside the model: incoherent videos consistently exhibit irregular, fragmented temporal diagonals in …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding

Ziyang Wang, Yue Zhang, Shoubin Yu, Ce Zhang et autres

Next-generation visual assistants, such as smart glasses, embodied agents, and always-on life-logging systems, must reason over an entire day or more of continuous visual experience. In ultra-long videos, relevant information is sparsely distributed across hours or days, making memory a fundamental challenge: …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation

Yan-Bo Lin, Jonah Casebeer, Long Mai, Aniruddha Mahapatra et autres

Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal control. We introduce V2M-ZERO, a video-to-music generation approach that generates time-aligned music with disentangled time synchronization and semantic control (e.g., genre, mood) from video …

0 citations arXiv (Cornell University)
Accès ouvert 2026 conference-paper OpenAlex

TimeRefine: Temporal Grounding with Time Refining Video LLM

Xizi Wang, Feng Cheng, Ziyang Wang, Huiyu Wang et autres

Video temporal grounding aims to localize relevant temporal boundaries in a video given a textual prompt. Recent work has focused on enabling Video LLMs to perform video temporal grounding via next-token prediction of temporal timestamps. However, accurately localizing timestamps in videos remains …

us, mx (code pays fourni par la source)

2 citations
Accès ouvert 2026 conference-paper OpenAlex

Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising

Yan-Bo Lin, Kevin Lin, Zhengyuan Yang, Linjie Li et autres

In this paper, we introduce zero-shot audio-video editing, a novel task that requires transforming original audio-visual content to align with a specified textual prompt without additional model training. To evaluate this task, we curate a benchmark dataset, AVED-Bench, designed explicitly for zero-shot …

us, gb (code pays fourni par la source)

0 citations
Accès ouvert 2026 conference-paper OpenAlex

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction

Ce Zhang, Yale Song, Ruta Desai, Michael L. Iuzzolino et autres

Visual Planning for Assistance (VPA) aims to predict a sequence of user actions required to achieve a specified goal based on a video showing the user’s progress. Although recent advances in multimodal large language models (MLLMs) have shown promising results in video …

us (code pays fourni par la source)

0 citations
Accès ouvert 2026 preprint OpenAlex

LiLo-VLA: Compositional Long-Horizon Manipulation via Linked Object-Centric Policies

Yue Yang, Shuo Cheng, Yu Fang, Homanga Bharadhwaj et autres

General-purpose robots must master long-horizon manipulation, defined as tasks involving multiple kinematic structure changes (e.g., attaching or detaching objects) in unstructured environments. While Vision-Language-Action (VLA) models offer the potential to master diverse atomic skills, they struggle with the combinatorial complexity of sequencing …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs

Baiqi Li, Kangyi Zhao, Ce Zhang, Chancharik Mitra et autres

Fine-grained spatio-temporal understanding is essential for video reasoning and embodied AI. Yet, while Multimodal Large Language Models (MLLMs) master static semantics, their grasp of temporal dynamics remains brittle. We present TimeBlind, a diagnostic benchmark for compositional spatio-temporal understanding. Inspired by cognitive science, …

1 citation arXiv (Cornell University)

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.