Coarse-to-Fine Multimodal Information Selection for Video Speaking Style Recognition with Large Language Models
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Video Speaking Style Recognition (VSSR) aims to classify conversation videos into different types, significantly facilitating human interaction understanding.Recent approaches explore the potential of large language models (LLM) in VSSR with a training-free process.However, directly integrating all multimodal data yields suboptimal results, since the great redundancy in visual data can overshadow other valuable multimodal information, such as valuable textual dialogues and critical visual clues.To address this, we propose CFMiS (Coarseto-Fine Multimodal Information Selection), a novel framework for VSSR that dynamically obtain valuable multimodal data via coarseto-fine selection, enhancing LLM reasoning for VSSR.Specifically, the core of CFMiS are two cascaded modules: 1) a text-dominant modality selection module firstly selects VSSRrequired modalities originating from text-based prediction; and 2) if vision is included in the selected modalities, a visual refinement module iteratively collects VSSR-relevant critical visual clues.The former resolves which modality to utilize, while the latter determines which information to adopt from selected modalities, efficiently alleviating information redundancy.Extensive experiments on multiple datasets prove that CFMiS is highly effective for VSSR, outperforming all existing trainingfree approaches and most training-based methods.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Coarse-to-Fine Multimodal Information Selection for Video Speaking Style Recognition with Large Language Models
- Date Crossref
- 01/01/2026
- Éditeur
- Association for Computational Linguistics
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Nanjing University State Key Laboratory for Novel Software Technology pays non établi dans la noticeUniversité ou école supérieure
-
Tencent (China) pays non établi dans la noticeEntreprise
State Key Laboratory for Novel Software Technology — Nanjing University et Tencent (China).
Une affiliation ne permet pas de déduire la nationalité d’un auteur.