Aller au contenu principal
Profil bibliographique

Zhucun Xue

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

26Publications signalées
950Citations signalées
1Affiliations récentes

Les institutions déclarées

Les domaines associés

COVID-19 diagnosis using AIGenerative Adversarial Networks and Image SynthesisDomain Adaptation and Few-Shot LearningAnomaly Detection Techniques and ApplicationsMultimodal Machine Learning Applications

Les publications récentes

Accès ouvert 2025 preprint OpenAlex

Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10$\times$

Jiangning Zhang, Junwei Zhu, Teng Hu, Yabiao Wang et autres

Native 4K (2160$\times$3840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, making it difficult for models to strike a balance between efficiency and quality. This paper proposes a novel Transformer retrofit strategy …

0 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

InstanceV: Instance-Level Video Generation

Yuheng Chen, Teng Hu, Jiangning Zhang, Zhucun Xue et autres

Recent advances in text-to-video diffusion models have enabled the generation of high-quality videos conditioned on textual descriptions. However, most existing text-to-video models rely solely on textual conditions, lacking general fine-grained controllability over video generation. To address this challenge, we propose InstanceV, a …

0 citations arXiv (Cornell University)
2025 conference-paper OpenAlex

Multi-Modal Retrieval Augmented Visual Understanding and Generation

Zhucun Xue

This paper presents a doctoral research focusing on integrating Retrieval-Augmented Generation (RAG) into video-related multimodal tasks. Existing RAG studies predominantly target text, images, or tabular data, overlooking the unique value of video as a knowledge carrier. We address this gap by: 1) …

cn (code pays fourni par la source)

1 citation
2025 article OpenAlex

ImitDiff: Transferring Foundation-Model Priors for Distraction-Robust Visuomotor Policy

Jiangning Zhang, Beiwen Tian, Hongrui Zhu, Yufei Jia et autres

Visuomotor imitation learning policies enable robots to efficiently acquire manipulation skills from visual demonstrations. However, as scene complexity and visual distractions increase, policies that perform well in simple settings often experience substantial performance degradation. To address this challenge, we propose ImitDiff, a …

cn (code pays fourni par la source)

1 citation IEEE Robotics and Automation Letters
2025 article OpenAlex

EMOv2: Pushing 5M Vision Model Frontier

Jiangning Zhang, Zhucun Xue, Yabiao Wang, Chengjie Wang et autres

This work focuses on developing parameter-efficient and lightweight models for dense predictions while trading off parameters, FLOPs, and performance. Our goal is to set up the new frontier of the 5 M magnitude lightweight model on various downstream tasks. Inverted Residual Block …

cn, sg (code pays fourni par la source)

5 citations IEEE Transactions on Pattern Analysis and Machine Intelligence
Accès ouvert 2025 preprint OpenAlex

UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions

Zhucun Xue, Jiangning Zhang, Teng Hu, Haoyang He et autres

The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. The growing demand for video applications sets higher requirements for high-quality video generation models. For example, the generation of movie-level Ultra-High …

0 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding

Zhucun Xue, Jiangning Zhang, Xurong Xie, Yuxuan Cai et autres

Multimodal Large Language Models (MLLMs) perform well in video understanding but degrade on long videos due to fixed-length context and weak long-term dependency modeling. Retrieval-Augmented Generation (RAG) can expand knowledge dynamically, yet existing video RAG schemes adopt fixed retrieval paradigms that ignore …

0 citations arXiv (Cornell University)
2025 conference-paper OpenAlex

Improving Autoregressive Visual Generation with Cluster-Oriented Token Prediction

Teng Hu, Jiangning Zhang, Ran Yi, Jieyu Weng et autres

Employing LLMs for visual generation has recently become a research focus. However, the existing methods primarily transfer the LLM architecture to visual generation but rarely investigate the fundamental differences between language and vision. This oversight may lead to suboptimal utilization of visual …

cn (code pays fourni par la source)

3 citations
2025 conference-paper OpenAlex

GroundingFace: Fine-grained Face Understanding via Pixel Grounding Multimodal Large Language Model

Yue Han, Jiangning Zhang, Junwei Zhu, Runze Hou et autres

Multimodal Language Learning Models (MLLMs) have shown remarkable performance in image understanding, generation, and editing, with recent advancements achieving pixel-level grounding with reasoning. However, these models for common objects struggle with fine-grained face understanding. In this work, we introduce the FacePlayGround-240K dataset, …

cn (code pays fourni par la source)

1 citation
2025 conference-paper OpenAlex

TIMotion: Temporal and Interactive Framework for Efficient Human-Human Motion Generation

Yabiao Wang, Shuo Wang, Jiangning Zhang, Ke Fan et autres

Human-human motion generation is essential for understanding humans as social beings. Current methods fall into two main categories: single-person-based methods and separate modeling-based methods. To delve into this field, we abstract the overall generation process into a general framework MetaMotion, which consists …

cn (code pays fourni par la source)

10 citations

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.