Enhancing Unsupervised Acoustic Word Embedding with Visual-Grounded Speech Model and Novel Word-level ABX Evaluation Schemes
Résumé fourni par la source
Most recent Acoustic Word Embedding (AWE) systems utilize an autoencoder-like approach to compress speech features of arbitrary shapes into fixed-size numerical vectors and then reconstructing it, thereby capturing essential patterns in the data. Unfortunately, AWE models have commonly relied on supervised learning, necessitating extensive textual data, or have employed unsupervised dynamic-based methods that are computationally demanding. This paper introduces an unsupervised approach to AWE, leveraging a self-supervised Visual-Grounded Speech (VGS) model, eliminating the need for dynamic algorithms or textual data. Additionally, we propose a fine-grained ABX evaluation protocol that meticulously assesses the acoustic similarity between spoken segments, providing a more comprehensive and fair evaluation of model performance. Our findings indicate that proposed visual-grounded approach allows the AWE model to function in a truly unsupervised manner without relying on text data and computationally intensive dynamic-based algorithms, while also achieving performance comparable to other approaches.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Enhancing Unsupervised Acoustic Word Embedding with Visual-Grounded Speech Model and Novel Word-level ABX Evaluation Schemes
- Date Crossref
- 06/04/2025
- Éditeur
- IEEE
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.