Aller au contenu principal
Profil bibliographique

Yi Xu

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

9Publications signalées
0Citations signalées
0Affiliations récentes

Les domaines associés

Robot Manipulation and LearningMultimodal Machine Learning ApplicationsReinforcement Learning in RoboticsGenerative Adversarial Networks and Image SynthesisHuman Pose and Action Recognition

Les publications récentes

Accès ouvert 2026 preprint OpenAlex

MaskHarness-WAM: Instance-Grounded Harnessing for Long-Horizon Robot Manipulation

Zitai Huang, Taiyi Su, Jian Zhu, Jianjun Zhang et autres

Long-horizon robot manipulation requires not only stable local visuomotor control, but also continuous target tracking and reliable task progress assessment throughout execution. This challenge becomes particularly critical when multiple objects share identical appearances and must be manipulated in a prescribed order. In …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models

Tianbin Liu, Jian Zhu, Taiyi Su, Jianjun Zhang et autres

FastWAM-style world action models enable efficient action-only inference, but generalize poorly under visual distribution shifts. Their reconstruction-oriented representations emphasize appearance-specific details, limiting generalization to unseen scenes and objects. Without observation history, the model also lacks temporal evidence for robustly identifying task-relevant state …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

KnowVis: Knowledge-Centric Visual Summarization for Video Lectures

Yi Xu, Yifan Hou, Xiaoyu Zhang

Video lectures are valuable educational resources, but their dense and lengthy formats often overwhelm novice learners. This difficulty stems from a fundamental pedagogical mismatch: while videos deliver transient information linearly, human learning requires constructing interconnected cognitive networks, a task that induces severe …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

SPLG-Mamba: Structure-Preserving Local-Global Mamba Network for Salient Object Detection in Optical Remote Sensing Images

Yi Xu, Ruichao Hou, Tongwei Ren, Gangshan Wu

Salient object detection in optical remote sensing images (ORSI-SOD) requires dense predictions that preserve object completeness and structural continuity under complex backgrounds, scale variation, and irregular object shapes. Existing methods often localize salient regions, but their predictions may still suffer from structural …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

GlanceWAM: Sparse Test-Time Imagination for World-Action Models

Linhan Wang, Zijian An, Mingyuan Zhang, Chen Dai et autres

Video generative models provide rich physical priors for robot learning, yet existing world-action models (WAMs) face a fundamental trade-off: synchronous video generation at control rate is latency-prohibitive, while abandoning test-time visual imagination sacrifices task success. We show that visual imagination achieves both …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

JEPA-WAM: Connecting Generated Visual Instructions to World Action Models through JEPA Latent Representations

Tianbin Liu, Jian Zhu, Taiyi Su, Jianjun Zhang et autres

World Action Models (WAMs) have demonstrated strong robotic manipulation capabilities by augmenting pretrained video generative models with action experts. However, current WAMs still show limited instruction-following ability when conditioned solely on text instructions. We argue that this limitation stems in part from …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation

J ZHU, Jianjun Zhang, Taiyi Su, Tianbin Liu et autres

World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learning. Existing WAMs excel at physically grounded execution, but typically lack the explicit language-level planning interface in VLM-based VLAs …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Learning 4D Geometric Priors for Inference-Efficient World Action Models

Jianjun Zhang, J ZHU, Taiyi Su, C W et autres

World Action Models (WAMs) have shown strong potential for robotic manipulation by jointly modeling visual future dynamics and executable action sequences. However, existing video-action co-training methods primarily optimize appearance-oriented video latents, which may insufficiently capture the temporally evolving geometry required for precise …

0 citations arXiv (Cornell University)

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.