Developing a Robust Mispronunciation Detection by Data Augmentation Based on Automatic Phone Annotation
Rattachement africain : kr. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Limited non-native English speech data and accurate pro- nunciation labels results in data sparsity in Mispronunciation Detection and Diagnosis (MDD), hindering its performance. The purpose of this study is to investigate the impact of data augmentation using automatic phone annotation on the per- formance of MDD. We proposed to utilize Self-Supervised Learning (SSL)-based phone recognizers and Recognizer Output Voting Error Reduction(ROVER) to automatically generate phone annotations of the ESLTTS dataset, a corpus of non-native English speech that lacks phone-level annotations. We then evaluate the performance of MDD tasks. The experimental results indicate that the proposed methodology yields improvements in both Phone Error Rate (PER) and F1 score.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Developing a Robust Mispronunciation Detection by Data Augmentation Based on Automatic Phone Annotation
- Date Crossref
- 17/10/2024
- Éditeur
- IEEE
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Seoul National University Interdisciplinary Program in Cognitive Science pays non établi dans la noticeUniversité ou école supérieure
Interdisciplinary Program in Cognitive Science — Seoul National University.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.