2026
conference-paper
OpenAlex
C. Li, Wei Wang, Marvin Sach, Wangyou Zhang et autres
The ICASSP 2026 URGENT Challenge advances the series by focusing on universal speech enhancement (SE) systems that handle diverse distortions, domains, and input conditions. This overview paper details the challenge’s motivation, task definitions, datasets, baseline systems, evaluation protocols, and results. The challenge …
cn, de, jp, us
(code pays fourni par la source)
2026
conference-paper
OpenAlex
Marvin Sach, Yihui Fu, Kohei Saijo, Wangyou Zhang et autres
In speech quality estimation for speech enhancement (SE) systems, subjective listening tests so far are considered as the gold standard. This should be even more true considering the large influx of new generative or hybrid methods into the field, revealing issues of …
de, jp, cn, us
(code pays fourni par la source)
2026
conference-paper
OpenAlex
Alexander Polok, Dominik Klement, Samuele Cornell, Matthew Wiesner et autres
Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a major challenge. While some approaches achieve strong performance when fine-tuned on specific domains, few systems generalize well across out-of-domain datasets. Our prior work, Diarization-Conditioned Whisper (DiCoW), leverages speaker diarization outputs as conditioning …
cz, us
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Chang Li, Wangyou Zhang, Wei Wang, Robin Scheibler et autres
The vast majority of modern speech enhancement systems rely on data-driven neural network models. Conventionally, larger datasets are presumed to yield superior model performance, an observation empirically validated across numerous tasks in other domains. However, recent studies reveal diminishing returns when scaling …
cn, us, jp, de
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Yi-Xiang Wang, Chang Li, Wei Wang, Wangyou Zhang et autres
The Mean Opinion Score (MOS) is fundamental to speech quality assessment. However, its acquisition requires significant human annotation. Although deep neural network approaches, such as DNSMOS and UTMOS, have been developed to predict MOS to avoid this issue, they often suffer from …
cn, us, de, jp
(code pays fourni par la source)
Accès ouvert
2025
article
OpenAlex
Samuele Cornell, Christoph Boeddeker, Tae‐Jin Park, He Huang et autres
us, de, cn, it
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi, Satoru Fukayama et autres
Masked token prediction has emerged as a powerful pre-training objective across language, vision, and speech, offering the potential to unify these diverse modalities through a single pre-training task. However, its application for general audio understanding remains underexplored, with BEATs being the only …
us, jp
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Brian Yan, Injy Hamed, Shuichiro Shimizu, Vasista Sai Lodagala et autres
We present CS-FLEURS, a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages. CS-FLEURS consists of 4 test sets which cover in total 113 unique code-switched language pairs across 52 languages: 1) a 14 X-English language …
2025
conference-paper
OpenAlex
Zaid Sheikh, Shuichiro Shimizu, Siddhant Arora, Jiatong Shi et autres
2025
conference-paper
OpenAlex
Wangyou Zhang, Kohei Saijo, Samuele Cornell, Robin Scheibler et autres
2025
conference-paper
OpenAlex
Kohei Saijo, Wangyou Zhang, Samuele Cornell, Robin Scheibler et autres
Accès ouvert
2025
preprint
OpenAlex
Samuele Cornell, Christoph Boeddeker, Tae‐Jin Park, He Huang et autres
The CHiME-7 and 8 distant speech recognition (DASR) challenges focus on multi-channel, generalizable, joint automatic speech recognition (ASR) and diarization of conversational speech. With participation from 9 teams submitting 32 diverse systems, these challenges have contributed to state-of-the-art research in the field. …