Accès ouvert
2026
preprint
OpenAlex
Zhenglong Liu, Wangyou Zhang, Chenda Li, Yanmin Qian
Multi-channel speech enhancement (SE) systems exhibit superior performance over single-channel methods but are constrained to fixed microphone array configurations. This restricts their real-world deployment across devices with diverse array geometries. While recent array-agnostic SE methods address variable microphone numbers and permutations, they …
Accès ouvert
2026
preprint
OpenAlex
Qinzhe Hu, Chenda Li, Wangyou Zhang, Liu S et autres
Recent advances in speech separation (SS) have led to compact front-end models with small parameter sizes, yet their high computational cost remains a major barrier for deployment on edge devices. To address this, we propose TF-MoE, a sparse Mixture-of-Experts (MoE) framework that …
Accès ouvert
2026
preprint
OpenAlex
Masao Someki, Alexander Polok, Carlos Carvalho, Chyi-Jiunn Lin et autres
Recent speech research involves increasingly large datasets, complex models, and diverse experimental workflows. However, existing frameworks require substantial engineering effort to support such experiments. We present ESPnet3, a speech and audio research framework built on a modular system architecture with configuration-driven dataset …
2026
conference-paper
OpenAlex
C. Li, Wei Wang, Marvin Sach, Wangyou Zhang et autres
The ICASSP 2026 URGENT Challenge advances the series by focusing on universal speech enhancement (SE) systems that handle diverse distortions, domains, and input conditions. This overview paper details the challenge’s motivation, task definitions, datasets, baseline systems, evaluation protocols, and results. The challenge …
cn, de, jp, us
(code pays fourni par la source)
2026
conference-paper
OpenAlex
Marvin Sach, Yihui Fu, Kohei Saijo, Wangyou Zhang et autres
In speech quality estimation for speech enhancement (SE) systems, subjective listening tests so far are considered as the gold standard. This should be even more true considering the large influx of new generative or hybrid methods into the field, revealing issues of …
de, jp, cn, us
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Changhao Cheng, Wei Wang, Wangyou Zhang, Dongya Jia et autres
Continuous speech representations based on Variational Autoencoders (VAEs) have emerged as a promising alternative to traditional spectrogram or discrete token based features for speech generation and reconstruction. Recent research has tried to enrich the structural information in VAE latent representations by aligning …
Accès ouvert
2026
preprint
OpenAlex
Bing Han, Chushu Zhou, Yifan Yang, Wei Wang et autres
Bootstrap-based Self-Supervised Learning (SSL) has achieved remarkable progress in audio understanding. However, existing methods typically operate at a single level of granularity, limiting their ability to model the diverse temporal and spectral structures inherent in complex audio signals. Furthermore, bootstrapping representations from …
Accès ouvert
2026
preprint
OpenAlex
Wei Wang, Wangyou Zhang, Chenda Li, Jiahe Wang et autres
Automatic speech quality assessment has become increasingly important as modern speech generation systems continue to advance, while human listening tests remain costly, time-consuming, and difficult to scale. Most existing learning-based assessment models rely primarily on scarce human-annotated mean opinion score (MOS) data, …
Accès ouvert
2026
preprint
OpenAlex
Chenda Li, Wei Wang, Marvin Sach, Wangyou Zhang et autres
The ICASSP 2026 URGENT Challenge advances the series by focusing on universal speech enhancement (SE) systems that handle diverse distortions, domains, and input conditions. This overview paper details the challenge's motivation, task definitions, datasets, baseline systems, evaluation protocols, and results. The challenge …
Accès ouvert
2026
preprint
OpenAlex
Chenda Li, Wei Wang, Marvin Sach, Wangyou Zhang et autres
The ICASSP 2026 URGENT Challenge advances the series by focusing on universal speech enhancement (SE) systems that handle diverse distortions, domains, and input conditions. This overview paper details the challenge's motivation, task definitions, datasets, baseline systems, evaluation protocols, and results. The challenge …
cn, de, jp, us, it
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Chang Li, Wangyou Zhang, Wei Wang, Robin Scheibler et autres
The vast majority of modern speech enhancement systems rely on data-driven neural network models. Conventionally, larger datasets are presumed to yield superior model performance, an observation empirically validated across numerous tasks in other domains. However, recent studies reveal diminishing returns when scaling …
cn, us, jp, de
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Yi-Xiang Wang, Chang Li, Wei Wang, Wangyou Zhang et autres
The Mean Opinion Score (MOS) is fundamental to speech quality assessment. However, its acquisition requires significant human annotation. Although deep neural network approaches, such as DNSMOS and UTMOS, have been developed to predict MOS to avoid this issue, they often suffer from …
cn, us, de, jp
(code pays fourni par la source)