Aller au contenu principal
2026 article

Speaker-Conditioned U-Shaped Diarization With Speaker Extraction-Guided Enhancement

0Citations signalées, ce qui n’est pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : vn. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Speaker diarization demarcates speech segments by speaker, answering the question “who spoke when?”. Recently, a promising approach has emerged by integrating speaker diarization with speech separation or speaker extraction, which offers better generalization and requires significantly less pretraining data. Despite this progress, efficiently aligning speaker extraction with the inherent requirements of diarization remains a key challenge. In this paper, we introduce Speaker-conditioned U-shaped Diarization with Speaker Extraction-guided Enhancement (SUDx), a joint network where speaker diarization is the primary target and speaker extraction serves as a supportive auxiliary task. SUDx leverages the U-net architecture and exploits hierarchical speaker representation to enhance the performance of speaker diarization. In addition, SUDx does not rely on explicit speaker identity labels for supervision, allowing it to learn distinctive acoustic characteristics and adapt easily to real-recorded multi-speaker conversations. Furthermore, we introduce a novel inference strategy that effectively handles unknown number of speakers and reduces reliance on large-scale pretraining data. We show that SUDx outperforms competitive baselines for speaker diarization while maintaining high-quality speech extraction on the LibriMix dataset. We further assess our proposed approach and our novel strategy on the AMI and AISHELL-4 meeting corpora, experimental results indicate that our model achieved state-of-the-art performance with much less pretraining data.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Speaker-Conditioned U-Shaped Diarization With Speaker Extraction-Guided Enhancement
Date Crossref
01/01/2026
Éditeur
Institute of Electrical and Electronics Engineers (IEEE)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Hanoi University of Science and Technology pays non établi dans la notice
    Université ou école supérieure
  • School of Electrical and Electronic Engineering pays non établi dans la notice
    Université ou école supérieure
  • Viettel AI pays non établi dans la notice
    Institution

Hanoi University of Science and Technology, School of Electrical and Electronic Engineering et Viettel AI.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Speech Recognition and SynthesisSpeech and Audio ProcessingEmotion and Mood Recognition

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.