Accès ouvert
2025
preprint
OpenAlex
Yuyang Hu, Shichun Liu, Yanwei Yue, Guibin Zhang et autres
Memory has emerged, and will continue to remain, a core capability of foundation model-based agents. As research on agent memory rapidly expands and attracts unprecedented attention, the field has also become increasingly fragmented. Existing works that fall under the umbrella of agent …
Accès ouvert
2025
preprint
OpenAlex
Hengfeng Gong, Xiaozhong Ji, Yuansen Liu, Wenbin Wu et autres
MLLMs MLLMs are beginning to appear in clinical workflows, but their ability to perform complex medical reasoning remains unclear. We present Med-CMR, a fine-grained Medical Complex Multimodal Reasoning benchmark. Med-CMR distinguishes from existing counterparts by three core features: 1) Systematic capability decomposition, …
Accès ouvert
2025
preprint
OpenAlex
Huayang Huang, Peng Xu, Xiaobin Hu, Donghao Luo et autres
Recent hybrid video generation models combine autoregressive temporal dynamics with diffusion-based spatial denoising, but their sequential, iterative nature leads to error accumulation and long inference times. In this work, we propose a distillation-based framework for efficient causal video generation that enables high-quality …
Accès ouvert
2025
conference-paper
OpenAlex
Weipeng Tan, Chuming Lin, Chengming Xu, FeiFan Xu et autres
Recent advances in Talking Head Generation (THG) have achieved impressive lip synchronization and visual quality through diffusion models; yet existing methods struggle to generate emotionally expressive portraits while preserving speaker identity. We identify three critical limitations in current emotional talking head generation: …
cn
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Yuji Wang, Moran Li, Xiaobin Hu, Ran Yi et autres
Identity-preserving text-to-video (IPT2V) generation, which aims to create high-fidelity videos with consistent human identity, has become crucial for downstream applications. However, current end-to-end frameworks suffer a critical spatial-temporal trade-off: optimizing for spatially coherent layouts of key elements ( e.g., character identity preservation) …
cn
(code pays fourni par la source)
2025
article
OpenAlex
Mingyu Liu, Jiong Xu, Yuning Cui, Xiaobin Hu et autres
Real-world image quality is often degraded by suboptimal lighting conditions, such as low light and vignetting. With advancements in sensor technology, there is an increasing demand for efficient algorithms capable of processing high-resolution images. However, existing approaches primarily focus on low-resolution data …
in, de, cn
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Caoshuo Li, Tanzhe Li, Xiaobin Hu, Donghao Luo et autres
Recently, Vision Graph Neural Network (ViG) has gained considerable attention in computer vision. Despite its groundbreaking innovation, Vision Graph Neural Network encounters key issues including the quadratic computational complexity caused by its K-Nearest Neighbor (KNN) graph construction and the limitation of pairwise …
cn
(code pays fourni par la source)
2025
article
OpenAlex
Peng Tang, Xiaoxiao Yan, Xiaobin Hu, Kai Wu et autres
Anomaly detection (AD) in medical applications is a promising field, offering a cost-effective alternative to labor-intensive abnormal data collection and labeling. However, the success of feature reconstruction-based methods in AD is often hindered by two critical factors: the domain gap of pre-trained …
de, cn, ch
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Weipeng Tan, Chuming Lin, Chengming Xu, FeiFan Xu et autres
Recent advances in Talking Head Generation (THG) have achieved impressive lip synchronization and visual quality through diffusion models; yet existing methods struggle to generate emotionally expressive portraits while preserving speaker identity. We identify three critical limitations in current emotional talking head generation: …
Accès ouvert
2025
preprint
OpenAlex
Caoshuo Li, Tanzhe Li, Xiaobin Hu, Donghao Luo et autres
Recently, Vision Graph Neural Network (ViG) has gained considerable attention in computer vision. Despite its groundbreaking innovation, Vision Graph Neural Network encounters key issues including the quadratic computational complexity caused by its K-Nearest Neighbor (KNN) graph construction and the limitation of pairwise …
2024
conference-paper
OpenAlex
Xiaobin Lu, Xiaobin Hu, Jun Luo, Yaping Ruan et autres
Blind face restoration endeavors to restore a clear face image from a degraded counterpart. Recent approaches employing Generative Adversarial Networks (GANs) as priors have demonstrated remarkable success in this field. However, these methods encounter challenges in achieving a balance between realism and …
cn
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Xiaobin Lu, Xiaobin Hu, Jun Luo, Ben Zhu et autres
Blind face restoration endeavors to restore a clear face image from a degraded counterpart. Recent approaches employing Generative Adversarial Networks (GANs) as priors have demonstrated remarkable success in this field. However, these methods encounter challenges in achieving a balance between realism and …