Accès ouvert
2026
preprint
OpenAlex
Junqi Liu, Yufan He, Yexiao He, Pengfei Guo et autres
Long-horizon agents are beginning to automate complete workflows that produce code, reports, and research artifacts. Medical imaging workflows are multi-stage and data-sensitive, while expert trajectories remain scarce and difficult to share. Structured benchmarks can localize failures through stage-level rubrics, but standard post-training …
Accès ouvert
2026
preprint
OpenAlex
Sreyan Ghosh, Arushi Goel, K. S. Jayakumar, Lasha Koroshinadze et autres
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over …
Accès ouvert
2026
preprint
OpenAlex
Sreyan Ghosh, Arushi Goel, K. S. Jayakumar, Lasha Koroshinadze et autres
We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over …
gb, us
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Shufan Li, Greg Heinrich, Hanrong Ye, Yonggan Fu et autres
We propose Nemotron-Labs-Diffusion-Image, a state-of-the-art masked discrete diffusion model (MDM) for high-resolution text-to-image synthesis. Compared with prior work on masked image generation, Nemotron-Labs-Diffusion-Image addresses two key challenges. First, unlike continuous diffusion models which progressively refine latent representations across the entire image, standard …
Accès ouvert
2026
preprint
OpenAlex
Shufan Li, Greg Heinrich, Hanrong Ye, Yonggan Fu et autres
We propose Nemotron-Labs-Diffusion-Image, a state-of-the-art masked discrete diffusion model (MDM) for high-resolution text-to-image synthesis. Compared with prior work on masked image generation, Nemotron-Labs-Diffusion-Image addresses two key challenges. First, unlike continuous diffusion models which progressively refine latent representations across the entire image, standard …
Accès ouvert
2026
preprint
OpenAlex
Johan Björck, Zhiqi Li, Yunze Man, Jing Wang et autres
Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, and long-horizon planning. While specialist models excel at individual tasks, deploying a multi-model stack is computationally expensive and prone to cascading errors. We present Vesta, a unified embodied generalist that …
Accès ouvert
2026
preprint
OpenAlex
Johan Björck, Zhiqi Li, Yunze Man, Yi-Xiang Wang et autres
Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, and long-horizon planning. While specialist models excel at individual tasks, deploying a multi-model stack is computationally expensive and prone to cascading errors. We present Vesta, a unified embodied generalist that …
us, hk, sg, de
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Yitong Jiang, Hongjun Wang, Collin McCarthy, Hanrong Ye et autres
Vision foundation models are bottlenecked by the quadratic cost of self-attention, which limits usable resolution and increases the cost of large-scale pretraining. Subquadratic alternatives such as linear attention and state-space models reduce this cost, but often serialize images into 1D token streams …
Accès ouvert
2026
preprint
OpenAlex
Yitong Jiang, Hongjun Wang, Collin McCarthy, Hanrong Ye et autres
Vision foundation models are bottlenecked by the quadratic cost of self-attention, which limits usable resolution and increases the cost of large-scale pretraining. Subquadratic alternatives such as linear attention and state-space models reduce this cost, but often serialize images into 1D token streams …
hk, us
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Dongyun Zou, Zhuoyang Zhang, Junyu Chen, Wenkun He et autres
We introduce JetViT, a novel family of hybrid-architecture Vision Transformer (ViT) models that match the accuracy of state-of-the-art full-attention vision foundation models while achieving substantially higher inference efficiency on high-resolution images. At the core of our approach is Post-Training Attention Search, a …
Accès ouvert
2026
preprint
OpenAlex
Dongyun Zou, Zhuoyang Zhang, Junyu Chen, Wenkun He et autres
We introduce JetViT, a novel family of hybrid-architecture Vision Transformer (ViT) models that match the accuracy of state-of-the-art full-attention vision foundation models while achieving substantially higher inference efficiency on high-resolution images. At the core of our approach is Post-Training Attention Search, a …
Accès ouvert
2026
preprint
OpenAlex
NVIDIA, :, Amala Sanjay Deshmukh, Kateryna Chumachenko et autres
We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, images, and video. Nemotron 3 Nano Omni delivers consistent accuracy improvements over its predecessor, Nemotron Nano V2 VL, …