Improving Continuous Sign Language Recognition via Cross-Frame Interactions in Expanded Contextual Spaces
Résumé fourni par la source
Current continuous sign language recognition (CSLR) methods typically rely on single or adjacent frames for calculations, which can overlook broader contextual information and result in lower accuracy. To address this issue, we introduce CVSign, which constructs an extended contextual space frame by frame while enabling comprehensive cross-frame interaction. Specifically, we present two innovative modules: Contextual Correspondence Awareness (CCA) and Contextual Variability Awareness (CVA). CCA enhances the relevance of contextual features by utilizing cross-frame multi-head query attention to identify and prioritize related areas while suppressing irrelevant regions. CVA captures motion changes at varying speeds by employing difference calculations between multiple frames, effectively minimizing static redundancy. Remarkably, experimental results show that CVSign outperforms the previous state-of-the-art method by a clear margin on widely used datasets, including PHOENIX14, PHOENIX14-T, and CSL-Daily.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Improving Continuous Sign Language Recognition via Cross-Frame Interactions in Expanded Contextual Spaces
- Date Crossref
- 06/04/2025
- Éditeur
- IEEE
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.