Accès ouvert
2025
preprint
OpenAlex
Brian Yan, Injy Hamed, Shuichiro Shimizu, Vasista Sai Lodagala et autres
We present CS-FLEURS, a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages. CS-FLEURS consists of 4 test sets which cover in total 113 unique code-switched language pairs across 52 languages: 1) a 14 X-English language …
2025
conference-paper
OpenAlex
Vasista Sai Lodagala, Lamya Alkanhal, Daniel Izham, Shivam Mehta et autres
2025
conference-paper
OpenAlex
Brian Yan, Injy Hamed, Shuichiro Shimizu, Vasista Sai Lodagala et autres
Accès ouvert
2025
conference-paper
OpenAlex
Sara Althubaiti, Vasista Sai Lodagala, Yousseif Ahmed Elshahawy, Daniel Izham et autres
Sara Althubaiti, Vasista Sai Lodagala, Tjad Clark, Yousseif Ahmed Elshahawy, Daniel Izham, Abdullah Alrajeh, Aljawahrah Bin Tamran, Ahmed Ali. Proceedings of The Third Arabic Natural Language Processing Conference. 2025.
2024
conference-paper
OpenAlex
Vasista Sai Lodagala, Abhishek Ghose Biswas, Shoutrik Das, Frank Jordan et autres
2023
conference-paper
OpenAlex
R Sivaguru, Vasista Sai Lodagala, S Umesh
Accès ouvert
2023
preprint
OpenAlex
R Sivaguru, Vasista Sai Lodagala, Srinivasan Umesh
While FastSpeech2 aims to integrate aspects of speech such as pitch, energy, and duration as conditional inputs, it still leaves scope for richer representations. As a part of this work, we leverage representations from various Self-Supervised Learning (SSL) models to enhance the …
2023
conference-paper
OpenAlex
Vasista Sai Lodagala, Sreyan Ghosh, Srinivasan Umesh
In this paper, we propose a new Self-Supervised Learning (SSL) algorithm called data2vec-aqc, for speech representation learning from unlabeled speech data. Our goal is to improve SSL for speech in domains where both unlabeled and labeled data are limited. Building on the …
in, us
(code pays fourni par la source)
2023
conference-paper
OpenAlex
Vasista Sai Lodagala, Sreyan Ghosh, Srinivasan Umesh
While Self-Supervised Learning has helped reap the benefit of the scale from the available unlabeled data, the learning paradigms are continously being bettered. We present a new pre-training strategy named ccc-wav2vec 2.0, which uses clustering and an augmentation based cross-contrastive loss as …
in, us
(code pays fourni par la source)
2023
conference-paper
OpenAlex
Vasista Sai Lodagala, Sreyan Ghosh, Srinivasan Umesh
While self-supervised speech representation learning (SSL) models serve a variety of downstream tasks, these models have been observed to overfit to the domain from which the unlabeled data originates. To alleviate this issue, we propose PADA (Pruning Assisted Domain Adaptation). Before performing …
in, us
(code pays fourni par la source)
Accès ouvert
2022
preprint
OpenAlex
Vasista Sai Lodagala, Sreyan Ghosh, S. Umesh
In this paper, we propose a new Self-Supervised Learning (SSL) algorithm called data2vec-aqc, for speech representation learning from unlabeled speech data. Our goal is to improve SSL for speech in domains where both unlabeled and labeled data are limited. Building on the …
Accès ouvert
2022
preprint
OpenAlex
Anusha Prakash, Arun Kumar, Ashish Seth, Bhagyashree Mukherjee et autres
Cross-lingual dubbing of lecture videos requires the transcription of the original audio, correction and removal of disfluencies, domain term discovery, text-to-text translation into the target language, chunking of text using target language rhythm, text-to-speech synthesis followed by isochronous lipsyncing to the original …