Phoneme-guided TTS augmentation for ASR: A unified pipeline and multilingual evaluation
Zhen Wang, T. Wu, 韩荣旗, Hao Wu et autres
Synthetic speech can provide additional supervision for automatic speech recognition (ASR), but constructing useful synthetic training data requires choosing both what to synthesize and how to synthesize it. We present a phoneme-guided text-to-speech (TTS) augmentation pipeline for ASR that connects multilingual speech …