GVEV-Net: Graph-based Visual Enhanced Video Network for Audio-Visual Question Answering
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Audio-visual question answering (AVQA) task, which aims to answer questions derived from the original videos, has attracted extensive attention in the fields of multimedia, computer vision, and natural language processing. Although the current studies have achieved promising results, there are still some limitations that need to be further explored. First, most existing methods mainly focus on frame-level visual features while paying little attention to object-level visual features, leading to ambiguous semantic representations of visual objects. Second, these methods ignore the temporal information embedded in the dynamic changes among adjacent time frames, causing the model to be unable to capture the temporal dependence in the video, thereby limiting the model’s capabilities. To address the challenges above, we propose the Graph-based Visual Enhanced Video Network (GVEV-Net), which enables fine-grained question-related visual information queries and guides the fusion of multimodal information through language information. Specifically, we first consider extracting audio-visual cues from object-level, frame-level visual information, and audio information. Then, we highlight the relevant object-level, and frame-level visual information and audio clips by taking the question as the guiding information. Next, graph attention is employed to capture the dynamic change relationship among adjacent frames in the object-level visual information. Finally, question-guided attention is used to fuse all the modal information. Experiments on the MUSIC-AVQA dataset demonstrate that our method outperforms seven existing recent methods. Moreover, the ablation study verifies the necessity and effectiveness of the GVEV-Net proposed in our proposal.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- GVEV-Net: Graph-based Visual Enhanced Video Network for Audio-Visual Question Answering
- Date Crossref
- 27/12/2024
- Éditeur
- IEEE
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Southwest University of Science and Technology pays non établi dans la noticeUniversité ou école supérieure
-
University of Science and Technology of China pays non établi dans la noticeUniversité ou école supérieure
-
Sichuan University of Arts and Science pays non établi dans la noticeUniversité ou école supérieure
-
Sichuan University of Culture and Arts pays non établi dans la noticeUniversité ou école supérieure
Southwest University of Science and Technology, University of Science and Technology of China et Sichuan University of Arts and Science, avec 1 autre affiliation.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.