Aller au contenu principal
2026 conference-paper

2025 Urgent Speech Enhancement Challenge Multilingual P.808 Listening Tests: Approach and Results

0Citations signalées, ce qui n’est pas une note de qualité
5Institutions déclarées
4Pays d’affiliation déclarés

Rattachement africain : de, jp, cn, us. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

In speech quality estimation for speech enhancement (SE) systems, subjective listening tests so far are considered as the gold standard. This should be even more true considering the large influx of new generative or hybrid methods into the field, revealing issues of some objective metrics. Efforts such as the Interspeech 2025 URGENT Speech Enhancement Challenge also involving non-English datasets add the aspect of multilinguality to the testing procedure. In this paper, we provide updated challenge results on the full multilingual test set and surprising insights into URGENT Challenge results, questioning the reliability of (P.808) absolute category rating (ACR) subjective testing as gold standard in the age of generative AI. Particularly, it seems that for generative SE methods, subjective (ACR MOS) and objective (DNSMOS, NISQA) reference-free metrics should be accompanied by objective phone fidelity metrics to reliably detect hallucinations. We briefly recap the ITU-T P.808 crowdsourced listening test method and detail our proposed and employed localization protocol for text and audio in subjective ACR tests, that we open-source, enabling reproducible evaluation of multilingual speech processing tasks.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
2025 Urgent Speech Enhancement Challenge Multilingual P.808 Listening Tests: Approach and Results
Date Crossref
03/05/2026
Éditeur
IEEE
Type
proceedings-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Technische Universität Braunschweig pays non établi dans la notice
    Université ou école supérieure
  • Waseda University pays non établi dans la notice
    Université ou école supérieure
  • Shanghai Jiao Tong University pays non établi dans la notice
    Université ou école supérieure
  • Carnegie Mellon University pays non établi dans la notice
    Université ou école supérieure
  • Google (United States) pays non établi dans la notice
    Entreprise
  • Technische Universit&#x00E4 pays non établi dans la notice
    Institution
  • Google DeepMind pays non établi dans la notice
    Institution
  • Meta pays non établi dans la notice
    Institution

Technische Universität Braunschweig, Waseda University et Shanghai Jiao Tong University, avec 5 autres affiliations.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Speech and Audio ProcessingHearing Loss and RehabilitationStuttering Research and Treatment

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.