Accès ouvert
2026
preprint
OpenAlex
Zhikun Zhou, Kunyu Peng, Runyi Yang, Junhao Cai et autres
Referring scene understanding for embodied robots requires grounding object- and relation-centric language queries from a designated viewpoint. While a local semantic Gaussian map can support such grounding within one agent's observations, cooperative settings require this ability to remain effective after independently reconstructed …
Accès ouvert
2026
preprint
OpenAlex
Carmelo Scribano, Filippo Muzzini, Nedyalko Prisadnikov, Mohammad Mahdi et autres
Beyond model inference, the decoding stage, which converts raw network outputs into task-level representations, constitutes a significant portion of the execution cost. Despite its practical impact, prediction decoding has received comparatively little attention and is often implemented using generic CPU routines or …
Accès ouvert
2026
preprint
OpenAlex
Anubhav Khanal, Prabigya Acharya, Roshni Poudel, Sujan Kapali et autres
Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current benchmarks. First, 3D datasets often rely on point clouds that capture geometry but discard rich visual features like texture, text, and …
Accès ouvert
2026
preprint
OpenAlex
Deheng Zhang, Letian Shi, Runyi Yang, Zhendong Li et autres
In rendering, a camera is a sampling operator that maps each finite pixel to a bundle of rays. Different camera models change the geometry of this bundle, thus making a unified and faithful rendering formulation challenging. Consequently, Gaussian ray tracing supports generic …
Accès ouvert
2026
preprint
OpenAlex
Runyi Yang, Deheng Zhang, Xiaoye Wang, Mengjiao Ma et autres
Language Gaussian fields implicitly assume that the primitive carrying semantics remains identifiable across views. This assumption breaks in scalable anchor-decoded representations, where persistent anchors generate view-conditioned child Gaussians whose geometry and appearance vary with the camera. We introduce Ours, a persistent language …
Accès ouvert
2026
preprint
OpenAlex
Ruibo Ming, Lei Sun, Deheng Zhang, He Zhang et autres
Many emerging video language modeling tasks require systems to move beyond clip-level abstraction and model visual content as it unfolds over extended time horizons. However, most existing video datasets rely on coarse or sparsely aligned supervision, which compresses temporal variation and limits …
Accès ouvert
2026
preprint
OpenAlex
Lei Sun, Yuqin Ma, Weilun Li, Haoran Liang et autres
Event-based motion estimation is central to tasks that demand high temporal resolution and robustness to fast motion. Existing methods typically rely on iterative optimization or repeated hypothesis comparison, offsetting the sensor's low-latency advantage. We propose Oriented Distance Field Motion Estimation (ODF Motion …
Accès ouvert
2026
preprint
OpenAlex
Sonali Godavarthy, Matthias Neuwirth-Trapp, Tim-Felix Faasch, Maarten Bieshaar et autres
Imaging factor disentanglement in text-to-image generation aims to independently control image acquisition properties such as types of camera lenses, sensor types, viewpoints, and domains to enable combinatorial generalization. This should let the model synthesize novel factor combinations unobserved in the training data, …
Accès ouvert
2026
preprint
OpenAlex
Saugat Adhikari, Ashok Prasad Neupane, Pramish Paudel, Ajad Chhatkuli et autres
Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints. This mismatch limits the usefulness of synthetic data for downstream training, where …
2026
conference-paper
OpenAlex
Sonali Godavarthy, Matthias Neuwirth-Trapp, Tim-Felix Faasch, Maarten Bieshaar et autres
de, ch, bg
(code pays fourni par la source)
2026
conference-paper
OpenAlex
Carmelo Scribano, Mohammad Mahdi, Nedyalko Prisadnikov, Yuqian Fu et autres
it, bg
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Stefan Maria Ailuro, Mario Markov, Mohammad Mahdi, Luc Van Gool et autres
Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this area has been driven by remote-sensing-specific architectural designs, often introducing new encoders, alignment modules, or task-specific fusion …