Accès ouvert
2026
preprint
OpenAlex
John Page, Xuesong Niu, Kai Wu, Kun Gai
Latent Diffusion Models (LDMs) rely heavily on the compressed latent space provided by Variational Autoencoders (VAEs) for high-quality image generation. Recent studies have attempted to obtain generation-friendly VAEs by directly adopting alignment strategies from LDM training, leveraging Vision Foundation Models (VFMs) as …
Accès ouvert
2026
preprint
OpenAlex
John Page, Xuesong Niu, Kai Wu, Kun Gai
Latent Diffusion Models (LDMs) rely heavily on the compressed latent space provided by Variational Autoencoders (VAEs) for high-quality image generation. Recent studies have attempted to obtain generation-friendly VAEs by directly adopting alignment strategies from LDM training, leveraging Vision Foundation Models (VFMs) as …
cn
(code pays fourni par la source)
Accès ouvert
2025
article
OpenAlex
Yong Li, Yi Ren, Xuesong Niu, Yi Ding et autres
Facial Action Units (AUs) are essential for conveying psychological states and emotional expressions. While automatic AU detection systems leveraging deep learning have progressed, they often overfit to specific datasets and individual features, limiting their cross-domain applicability. To overcome these limitations, we propose …
cn, sg
(code pays fourni par la source)
2024
conference-paper
OpenAlex
Xuesong Niu, Nan Jiang, Ruimao Zhang, Siyuan Huang
cn
(code pays fourni par la source)
2024
conference-paper
OpenAlex
Baoxiong Jia, Yixin Chen, Huangyue Yu, Yan Wang et autres
cn
(code pays fourni par la source)
2024
conference-paper
OpenAlex
Ziyu Zhu, Xiaojian Ma, Xuesong Niu, Yixin Chen et autres
cn
(code pays fourni par la source)
2024
article
OpenAlex
Xin Liu, Kaishen Yuan, Xuesong Niu, Jingang Shi et autres
Facial Action Unit (AU) detection is a crucial task in affective computing and social robotics as it helps to identify emotions expressed through facial expressions. Anatomically, there are innumerable correlations between AUs, which contain rich information and are vital for AU detection. …
cn
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Xiongkun Linghu, Jiangyong Huang, Xuesong Niu, Xiaojian Ma et autres
Situation awareness is essential for understanding and reasoning about 3D scenes in embodied AI agents. However, existing datasets and benchmarks for situated understanding are limited in data modality, diversity, scale, and task scope. To address these limitations, we propose Multi-modal Situated Question …
Accès ouvert
2024
preprint
OpenAlex
Ziyu Zhu, Zhuofan Zhang, Xiaojian Ma, Xuesong Niu et autres
A unified model for 3D vision-language (3D-VL) understanding is expected to take various scene representations and perform a wide range of tasks in a 3D scene. However, a considerable gap exists between existing methods and such a unified model, due to the …
Accès ouvert
2024
preprint
OpenAlex
Baoxiong Jia, Yixin Chen, Huangyue Yu, Yan Wang et autres
3D vision-language grounding, which focuses on aligning language with the 3D physical environment, stands as a cornerstone in the development of embodied agents. In comparison to recent advancements in the 2D domain, grounding language in 3D scenes faces several significant challenges: (i) …
Accès ouvert
2023
preprint
OpenAlex
Xin Liu, Kaishen Yuan, Xuesong Niu, Jingang Shi et autres
Facial Action Unit (AU) detection is a crucial task in affective computing and social robotics as it helps to identify emotions expressed through facial expressions. Anatomically, there are innumerable correlations between AUs, which contain rich information and are vital for AU detection. …
2023
conference-paper
OpenAlex
Hao Lu, Zitong Yu, Xuesong Niu, Ying-Cong Chen
Remote photoplethysmography (rPPG) technology has drawn increasing attention in recent years. It can extract Blood Volume Pulse (BVP) from facial videos, making many applications like health monitoring and emotional analysis more accessible. However, as the BVP signal is easily affected by environmental …
hk
(code pays fourni par la source)