Aller au contenu principal
Profil bibliographique

Xiaojian Ma

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

80Publications signalées
1531Citations signalées
1Affiliations récentes

Les institutions déclarées

Les domaines associés

Multimodal Machine Learning ApplicationsBiofuel production and bioconversionNatural Language Processing TechniquesHuman Pose and Action RecognitionCatalysis for Biomass Conversion

Les publications récentes

Accès ouvert 2025 conference-paper OpenAlex

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

Ziyu Zhu, Xilin Wang, Yixuan Li, Zhuofan Zhang et autres

Embodied scene understanding requires not only comprehending visual-spatial information that has been observed but also determining where to explore next in the 3D physical world. Existing 3D Vision-Language (3D-VL) models primarily focus on grounding objects in static observations from 3D reconstruction, such …

cn (code pays fourni par la source)

0 citations
2025 conference-paper OpenAlex

Adaptive Relay Hindsight Experience Replay for Sequential Manipulation Tasks with Sparse Rewards

Yuming Huang, Xiaojian Ma, Bin Fang, Huaping Liu et autres

In reinforcement learning, the sequential manipulation tasks with sparse rewards remain a great challenge. Recent approaches such as relay hindsight experience replay (RHER) expedite learning by reusing the already-learned policy with the self-guided exploration strategy (SGES). However, SGES causes decision conflicts between …

cn (code pays fourni par la source)

1 citation
2025 conference-paper OpenAlex

ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting

Shaofei Cai, Zihao Wang, Zhancun Mu, Xiaojian Ma et autres

Vision-language models (VLMs) have excelled in multimodal tasks, but adapting them to embodied decision-making in open-world environments presents challenges. One critical issue is bridging the gap between discrete entities in low-level observations and the abstract concepts required for effective planning. A common …

cn, us (code pays fourni par la source)

2 citations
Accès ouvert 2025 preprint OpenAlex

FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation

Jun Guo, Xiaojian Ma, Yi‐Kai Wang, Min Yang et autres

This paper investigates training better visual world models for robot manipulation, i.e., models that can predict future visual observations by conditioning on past frames and robot actions. Specifically, we consider world models that operate on RGB-D frames (RGB-D world models). As opposed …

0 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents

Bofei Zhang, Gao Zhi, Wang Zhang, Rui Xie et autres

Building Graphical User Interface (GUI) agents is a promising research direction, which simulates human interaction with computers or mobile phones to perform diverse GUI tasks. However, a major challenge in developing generalized GUI agents is the lack of sufficient trajectory data across …

0 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Muyao Li, Zihao Wang, Kaichen He, Xiaojian Ma et autres

Recently, action-based decision-making in open-world environments has gained significant attention. Visual Language Action (VLA) models, pretrained on large-scale web datasets, have shown promise in decision-making tasks. However, previous work has primarily focused on action post-training, often neglecting enhancements to the foundational model …

0 citations arXiv (Cornell University)
Accès ouvert 2025 conference-paper OpenAlex

JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Muyao Li, Zihao Wang, Kaichen He, Xiaojian Ma et autres

Recently, action-based decision-making in open-world environments has gained significant attention.Visual Language Action (VLA) models, pretrained on large-scale web datasets, have shown promise in decision-making tasks.However, previous work has primarily focused on action post-training, often neglecting enhancements to the foundation model itself.In response, …

cn (code pays fourni par la source)

4 citations
Accès ouvert 2024 preprint OpenAlex

Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding

Yue Fan, Xiaojian Ma, Rongpeng Su, Jun Guo et autres

This paper investigates the problem of understanding dynamic 3D scenes from egocentric observations, a key challenge in robotics and embodied AI. Unlike prior studies that explored this as long-form video understanding and utilized egocentric video only, we instead propose an LLM-based agent, …

0 citations arXiv (Cornell University)

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.