Aller au contenu principal
Profil bibliographique

Wangyou Zhang

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

115Publications signalées
1970Citations signalées
1Affiliations récentes

Les institutions déclarées

Les domaines associés

Speech Recognition and SynthesisSpeech and Audio ProcessingMusic and Audio ProcessingNatural Language Processing TechniquesAdvanced Adaptive Filtering Techniques

Les publications récentes

Accès ouvert 2026 preprint OpenAlex

Towards Array-Invariant Speech Enhancement via Geometry-Aware Dynamic Convolution

Zhenglong Liu, Wangyou Zhang, Chenda Li, Yanmin Qian

Multi-channel speech enhancement (SE) systems exhibit superior performance over single-channel methods but are constrained to fixed microphone array configurations. This restricts their real-world deployment across devices with diverse array geometries. While recent array-agnostic SE methods address variable microphone numbers and permutations, they …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era

Masao Someki, Alexander Polok, Carlos Carvalho, Chyi-Jiunn Lin et autres

Recent speech research involves increasingly large datasets, complex models, and diverse experimental workflows. However, existing frameworks require substantial engineering effort to support such experiments. We present ESPnet3, a speech and audio research framework built on a modular system architecture with configuration-driven dataset …

0 citations arXiv (Cornell University)
2026 conference-paper OpenAlex

ICASSP 2026 Urgent Speech Enhancement Challenge

C. Li, Wei Wang, Marvin Sach, Wangyou Zhang et autres

The ICASSP 2026 URGENT Challenge advances the series by focusing on universal speech enhancement (SE) systems that handle diverse distortions, domains, and input conditions. This overview paper details the challenge’s motivation, task definitions, datasets, baseline systems, evaluation protocols, and results. The challenge …

cn, de, jp, us (code pays fourni par la source)

2 citations
2026 conference-paper OpenAlex

2025 Urgent Speech Enhancement Challenge Multilingual P.808 Listening Tests: Approach and Results

Marvin Sach, Yihui Fu, Kohei Saijo, Wangyou Zhang et autres

In speech quality estimation for speech enhancement (SE) systems, subjective listening tests so far are considered as the gold standard. This should be even more true considering the large influx of new generative or hybrid methods into the field, revealing issues of …

de, jp, cn, us (code pays fourni par la source)

0 citations
Accès ouvert 2026 preprint OpenAlex

On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation

Changhao Cheng, Wei Wang, Wangyou Zhang, Dongya Jia et autres

Continuous speech representations based on Variational Autoencoders (VAEs) have emerged as a promising alternative to traditional spectrogram or discrete token based features for speech generation and reconstruction. Recent research has tried to enrich the structural information in VAE latent representations by aligning …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Representation-Regularized Convolutional Audio Transformer for Audio Understanding

Bing Han, Chushu Zhou, Yifan Yang, Wei Wang et autres

Bootstrap-based Self-Supervised Learning (SSL) has achieved remarkable progress in audio understanding. However, existing methods typically operate at a single level of granularity, limiting their ability to model the diverse temporal and spectral structures inherent in complex audio signals. Furthermore, bootstrapping representations from …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

UrgentMOS: Unified Multi-Metric and Preference Learning for Robust Speech Quality Assessment

Wei Wang, Wangyou Zhang, Chenda Li, Jiahe Wang et autres

Automatic speech quality assessment has become increasingly important as modern speech generation systems continue to advance, while human listening tests remain costly, time-consuming, and difficult to scale. Most existing learning-based assessment models rely primarily on scarce human-annotated mean opinion score (MOS) data, …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

ICASSP 2026 URGENT Speech Enhancement Challenge

Chenda Li, Wei Wang, Marvin Sach, Wangyou Zhang et autres

The ICASSP 2026 URGENT Challenge advances the series by focusing on universal speech enhancement (SE) systems that handle diverse distortions, domains, and input conditions. This overview paper details the challenge's motivation, task definitions, datasets, baseline systems, evaluation protocols, and results. The challenge …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

ICASSP 2026 URGENT Speech Enhancement Challenge

Chenda Li, Wei Wang, Marvin Sach, Wangyou Zhang et autres

The ICASSP 2026 URGENT Challenge advances the series by focusing on universal speech enhancement (SE) systems that handle diverse distortions, domains, and input conditions. This overview paper details the challenge's motivation, task definitions, datasets, baseline systems, evaluation protocols, and results. The challenge …

cn, de, jp, us, it (code pays fourni par la source)

0 citations arXiv (Cornell University)
2025 conference-paper OpenAlex

Less is More: Data Curation Matters in Scaling Speech Enhancement

Chang Li, Wangyou Zhang, Wei Wang, Robin Scheibler et autres

The vast majority of modern speech enhancement systems rely on data-driven neural network models. Conventionally, larger datasets are presumed to yield superior model performance, an observation empirically validated across numerous tasks in other domains. However, recent studies reveal diminishing returns when scaling …

cn, us, jp, de (code pays fourni par la source)

4 citations
2025 conference-paper OpenAlex

URGENT-PK: Perceptually-Aligned Ranking Model Designed for Speech Enhancement Competition

Yi-Xiang Wang, Chang Li, Wei Wang, Wangyou Zhang et autres

The Mean Opinion Score (MOS) is fundamental to speech quality assessment. However, its acquisition requires significant human annotation. Although deep neural network approaches, such as DNSMOS and UTMOS, have been developed to predict MOS to avoid this issue, they often suffer from …

cn, us, de, jp (code pays fourni par la source)

1 citation

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.