Aller au contenu principal
Profil bibliographique

Vasista Sai Lodagala

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

13Publications signalées
36Citations signalées
0Affiliations récentes

Les domaines associés

Speech Recognition and SynthesisNatural Language Processing TechniquesSpeech and dialogue systemsTopic ModelingSpeech and Audio Processing

Les publications récentes

Accès ouvert 2025 preprint OpenAlex

CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset

Brian Yan, Injy Hamed, Shuichiro Shimizu, Vasista Sai Lodagala et autres

We present CS-FLEURS, a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages. CS-FLEURS consists of 4 test sets which cover in total 113 unique code-switched language pairs across 52 languages: 1) a 14 X-English language …

0 citations arXiv (Cornell University)
Accès ouvert 2023 preprint OpenAlex

SALTTS: Leveraging Self-Supervised Speech Representations for improved Text-to-Speech Synthesis

R Sivaguru, Vasista Sai Lodagala, Srinivasan Umesh

While FastSpeech2 aims to integrate aspects of speech such as pitch, energy, and duration as conditional inputs, it still leaves scope for richer representations. As a part of this work, we leverage representations from various Self-Supervised Learning (SSL) models to enhance the …

0 citations arXiv (Cornell University)
2023 conference-paper OpenAlex

Data2vec-Aqc: Search for the Right Teaching Assistant in the Teacher-Student Training Setup

Vasista Sai Lodagala, Sreyan Ghosh, Srinivasan Umesh

In this paper, we propose a new Self-Supervised Learning (SSL) algorithm called data2vec-aqc, for speech representation learning from unlabeled speech data. Our goal is to improve SSL for speech in domains where both unlabeled and labeled data are limited. Building on the …

in, us (code pays fourni par la source)

2 citations
2023 conference-paper OpenAlex

CCC-WAV2VEC 2.0: Clustering AIDED Cross Contrastive Self-Supervised Learning of Speech Representations

Vasista Sai Lodagala, Sreyan Ghosh, Srinivasan Umesh

While Self-Supervised Learning has helped reap the benefit of the scale from the available unlabeled data, the learning paradigms are continously being bettered. We present a new pre-training strategy named ccc-wav2vec 2.0, which uses clustering and an augmentation based cross-contrastive loss as …

in, us (code pays fourni par la source)

12 citations 2022 IEEE Spoken Language Technology Workshop (SLT)
2023 conference-paper OpenAlex

PADA: Pruning Assisted Domain Adaptation for Self-Supervised Speech Representations

Vasista Sai Lodagala, Sreyan Ghosh, Srinivasan Umesh

While self-supervised speech representation learning (SSL) models serve a variety of downstream tasks, these models have been observed to overfit to the domain from which the unlabeled data originates. To alleviate this issue, we propose PADA (Pruning Assisted Domain Adaptation). Before performing …

in, us (code pays fourni par la source)

8 citations 2022 IEEE Spoken Language Technology Workshop (SLT)
Accès ouvert 2022 preprint OpenAlex

Technology Pipeline for Large Scale Cross-Lingual Dubbing of Lecture Videos into Multiple Indian Languages

Anusha Prakash, Arun Kumar, Ashish Seth, Bhagyashree Mukherjee et autres

Cross-lingual dubbing of lecture videos requires the transcription of the original audio, correction and removal of disfluencies, domain term discovery, text-to-text translation into the target language, chunking of text using target language rhythm, text-to-speech synthesis followed by isochronous lipsyncing to the original …

2 citations arXiv (Cornell University)

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.