Accès ouvert
2026
preprint
OpenAlex
Dongyoung Kim, Sumin Park, Woomin Song, Seungku Kim et autres
Improving embodied reasoning in multimodal-large-language models (MLLMs) is essential for building vision-language-action models (VLAs) on top of them to readily translate multimodal understanding into low-level actions. Accordingly, recent work has explored enhancing embodied reasoning in MLLMs through supervision of vision-question-answering type. However, …
Accès ouvert
2026
preprint
OpenAlex
Dongyoung Kim, Sumin Park, Woomin Song, Seungku Kim et autres
Improving embodied reasoning in multimodal-large-language models (MLLMs) is essential for building vision-language-action models (VLAs) on top of them to readily translate multimodal understanding into low-level actions. Accordingly, recent work has explored enhancing embodied reasoning in MLLMs through supervision of vision-question-answering type. However, …
Accès ouvert
2026
preprint
OpenAlex
Yuanhang Zhang, Younggyo Seo, Juyue Chen, Yifu Yuan et autres
Humanoid perceptive locomotion has made significant progress and shows great promise, yet achieving robust multi-directional locomotion on complex terrains remains underexplored. To tackle this challenge, we propose RPL, a two-stage training framework that enables multi-directional locomotion on challenging terrains, and remains robust …
Accès ouvert
2026
preprint
OpenAlex
Yuanhang Zhang, Younggyo Seo, Juyue Chen, Yifu Yuan et autres
Humanoid perceptive locomotion has made significant progress and shows great promise, yet achieving robust multi-directional locomotion on complex terrains remains underexplored. To tackle this challenge, we propose RPL, a two-stage training framework that enables multi-directional locomotion on challenging terrains, and remains robust …
Accès ouvert
2025
preprint
OpenAlex
Younggyo Seo, Carmelo Sferrazza, Guanya Shi, Rocky Duan et autres
Massively parallel simulation has reduced reinforcement learning (RL) training time for robots from days to minutes. However, achieving fast and reliable sim-to-real RL for humanoid control remains difficult due to the challenges introduced by factors such as high dimensionality and domain randomization. …
Accès ouvert
2025
preprint
OpenAlex
Huiwon Jang, Sihyun Yu, Heeseung Kwon, Younggyo Seo et autres
Leveraging temporal context is crucial for success in partially observable robotic tasks. However, prior work in behavior cloning has demonstrated inconsistent performance gains when using multi-frame observations. In this paper, we introduce ContextVLA, a policy model that robustly improves robotic task performance …
Accès ouvert
2025
preprint
OpenAlex
Tae‐Young Kim, Jimin Lee, M. Koo, Dongyoung Kim et autres
Vision-Language-Action (VLA) models have shown strong capabilities in robot manipulation by leveraging rich representations from pre-trained Vision-Language Models (VLMs). However, their representations arguably remain suboptimal, lacking sensitivity to robotic signals such as control actions and proprioceptive information. To address the issue, we …
Accès ouvert
2025
preprint
OpenAlex
Daewon Choi, Taeyoung Kim, Kyungmin Lee, Chang‐Yeon Kim et autres
Inherently, robotic manipulation tasks are history-dependent: leveraging past context could be beneficial. However, most existing Vision-Language-Action models (VLAs) have been designed without considering this aspect, i.e., they rely solely on the current observation, ignoring preceding context. In this paper, we propose HAMLET, …
2025
conference-paper
OpenAlex
Huiwon Jang, Sihyun Yu, Jinwoo Shin, Pieter Abbeel et autres
Efficient tokenization of videos remains a challenge in training vision models that can process long videos. One promising direction is to develop a tokenizer that can encode long video clips, as it would enable the tokenizer to leverage the temporal coherence of …
kr, ca, us
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Younggyo Seo, Carmelo Sferrazza, Haoran Geng, Michal Nauman et autres
Reinforcement learning (RL) has driven significant progress in robotics, but its complexity and long training times remain major bottlenecks. In this report, we introduce FastTD3, a simple, fast, and capable RL algorithm that significantly speeds up training for humanoid robots in popular …
Accès ouvert
2025
conference-paper
OpenAlex
Dongyoung Kim, Sumin Park, Huiwon Jang, Jinwoo Shin et autres
Large Vision-Language Models (LVLMs) have recently shown great promise in advancing robotics by combining embodied reasoning with robot control. A common approach involves training on embodied reasoning tasks related to robot control using Supervised Fine-Tuning (SFT). However, SFT datasets are often heuristically …
kr, ca, de
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Huiwon Jang, Sihyun Yu, Jinwoo Shin, Pieter Abbeel et autres
Efficient tokenization of videos remains a challenge in training vision models that can process long videos. One promising direction is to develop a tokenizer that can encode long video clips, as it would enable the tokenizer to leverage the temporal coherence of …