Accès ouvert
2026
conference-paper
OpenAlex
Jialong Mai, Jinxin Ji, Xiaofen Xing, Yang Chen et autres
Mainstream Automatic Speech Recognition (ASR) systems excel at transcribing lexical content, but largely fail to recognize nonverbal vocalizations (NVs) embedded in speech, such as sighs, laughs, and coughs. This capability is important for a comprehensive understanding of human communication, as NVs convey …
cn, hk
(code pays fourni par la source)
Accès ouvert
2026
conference-paper
OpenAlex
Sihang Nie, Xiaofen Xing, Jing Xing, Baiji Liu et autres
Large Language Model (LLM)-based Text-to-Speech (TTS) models have already reached a high degree of naturalness. However, the precision control of TTS inference is still challenging. Although instruction-based Text-to-Speech (Instruct-TTS) models are proposed, these models still lack fine-grained control due to the modality …
cn
(code pays fourni par la source)
2026
article
OpenAlex
Huachao Yan, Kailing Guo, Shiwei Song, Xiaofen Xing et autres
cn
(code pays fourni par la source)
Accès ouvert
2026
article
OpenAlex
Huimin Zheng, Xiaoqi Ai, Xueyan Liu, Xiaofen Xing et autres
Introduction: Enabling personalized sleep analysis and interaction directly on edge devices is crucial for providing real-time health insights and tailored guidance. However, this goal remains challenging due to the scarcity of high-quality physiological data and the computational constraints of edge hardware. Methods: …
cn
(code pays fourni par la source)
Accès ouvert
2026
conference-paper
OpenAlex
Wenyu Tao, Xiaofen Xing, Zeliang Li, Xiangmin Xu
2026
article
OpenAlex
Baoliang Feng, Lin Shu, Xiaofen Xing, Jun Guo et autres
cn
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Ziyi Zhao, Kailing Guo, Lin Wang, Fang Liu et autres
cn
(code pays fourni par la source)
2025
article
OpenAlex
Guodong Liang, Han Chen, Xiaofen Xing, Lan Zhang et autres
Abstract Objective. To develop a comprehensive physiological dataset for assessing internal and external stress and to propose robust automated stress recognition methods based on photoplethysmographic (PPG) signals. Approach. We established the Internal and External Stress Dataset (IESD), comprising PPG signals from 107 …
cn
(code pays fourni par la source)
2025
conference-paper
OpenAlex
P. L. Liu, Zhaojie Chu, Xiaofen Xing, Xiangmin Xu
Co-speech gestures are generally categorized into rhythmic and semantic gestures: rhythmic gestures align with speech rhythm and intonation, while semantic gestures convey specific meanings or emotions, enriching verbal communication. Most previous studies have focused on synthesizing rhythmic gestures, while recent methods have …
cn
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Sijin Yu, Zijiao Chen, Wenliang Wu, Shengxian Chen et autres
Reconstructing visual stimuli from human brain activity (e.g., fMRI) bridges neuroscience and computer vision by decoding neural representations. However, existing methods often overlook critical brain structure-function relationships, flattening spatial information and neglecting individual anatomical variations. To address these issues, we propose (1) …
cn, us
(code pays fourni par la source)
Accès ouvert
2025
article
OpenAlex
Yuyuan Chen, Cheng Xu, Xuemiao Xu, Xiaofen Xing et autres
Abstract 3D face swapping has been widely applied in entertainment, gaming and privacy protection. Traditional methods often interpolate directly in the latent space to generate a global swapped code, followed by per‐image latent code optimisation. While these approaches achieve 3D consistency, they …
hk, cn, sg
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Yongyi Su, Haojie Zhang, Shijie Li, Nanqing Liu et autres
Multimodal large language models (MLLMs) have advanced rapidly in recent years. However, existing approaches for vision tasks often rely on indirect representations, such as generating coordinates as text for detection, which limits performance and prevents dense prediction tasks like segmentation. To overcome …