SAQCodec: Semantic-Acoustic Fusion and Adaptive Quantization for Ultra-Low Bitrate Speech Coding
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Neural speech codecs have recently achieved remarkable progress in low-bitrate speech compression. However, effectively leveraging semantic information and addressing quantization limitations remain challenging, especially under ultra-low bitrate constraints. These factors may affect perceptual naturalness and reconstruction fidelity. To address these issues, we proposeSAQCodec, a neural speech codec that integrates a semantic-aware encoder, adaptive progressive quantization, and distribution-consistent training. Specifically, the semantic-aware encoder incorporates a Dual Cross-Attention Feature Fusion module to facilitate semantic-acoustic alignment. In addition, a Contextual Learning Module trained with a masked language modeling objective is introduced to capture global contextual dependencies. Furthermore, we present an Adaptive Progressive Residual Vector Quantization (APRVQ) scheme, which incorporates an attention weighting (AW) mechanism upon multi-scale RVQ to adaptively aligns quantization precision with feature importance, prioritizing critical features for finer quantization. A Conditional Entropy Quantization loss is employed as an auxiliary regularization to stabilize the quantization process. Experiments at ultra-low bitrates demonstrate that SAQCodec achieves clear performance advantages at extremely low bitrates (0.35 and 0.30 kbps), while remaining competitive with state-of-the-art codecs at higher ultra-low bitrate settings (0.97 and 0.47 kbps) in terms of reconstruction quality and downstream performance. Demos are available athttps://huazhi1024.github.io/SAQCodecDemo/.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- SAQCodec: Semantic-Acoustic Fusion and Adaptive Quantization for Ultra-Low Bitrate Speech Coding
- Date Crossref
- 01/01/2026
- Éditeur
- Institute of Electrical and Electronics Engineers (IEEE)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Wuhan University pays non établi dans la noticeUniversité ou école supérieure
-
School of Computer Science National Engineering Research Center for Multimedia Software (NERCMS) pays non établi dans la noticeUniversité ou école supérieure
Wuhan University et National Engineering Research Center for Multimedia Software (NERCMS) — School of Computer Science.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.