Towards Language-universal Mandarin-English Speech Recognition with Unsupervised Label Synchronous Adaptation
Rattachement africain : cn, us. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
End-to-end multilingual and code-switching speech recognition are two challenging tasks that are studied separately in many previous works. In this work, we jointly study multilingual and code-switching problems and present a novel unsupervised label synchronous adaptation algorithm for Mandarin-English speech recognition. Specifically, we use two parallel encoders to decompose the Mel-spectrum of speech into semantic information and other acoustic attributes, such as speaker identity, accents, pronunciation characteristics of different languages, etc. During the autoregressive decoding process of the speech recognition system, an adaptive decoder is used in parallel with the speech recognition decoder to generate an adaptive embedding for each character, so that the speech recognition model can be adaptive for Mandarin, English, and code-switching cases. Our experiments show that our proposed algorithm obtains 13.5% relative error reduction over a strong baseline in the code-switching case, and outperforms both the state-of-the-art Mandarin and English monolingual models.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Towards Language-universal Mandarin-English Speech Recognition with Unsupervised Label Synchronous Adaptation
- Date Crossref
- 11/12/2022
- Éditeur
- IEEE
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.