From Skeleton to Flesh: Aggregated Relational Transformer Towards Controllable Video Captioning with Two-Step Decoding
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Video captioning is the task of automatically producing natural language descriptions for a given video. The great success of the Transformer architecture in NLP and CV has inspired recent attempts at constructing an end-to-end Transformer-based model for video captioning. Although achieving impressive performance, the one-step framework relies on brute force to learn highly compressed semantics from a large number of video patches, lacking the review process, which is a common behavior of human beings for understanding videos/papers/books. Taking the rough but significant impression captured at first sight as the skeleton, additional review can verify and investigate specific information as flesh to deepen precise recognition. In this work, we introduce the review process in the Transformer-based framework and propose a novel network, Aggregated Relational Transformer (ART) to conduct two-step decoding for video captioning. Since the relation triplet concisely summarizes the main structure, it is assigned as the objective of the first-pass decoding. Then the relations are utilized as a skeleton and guide the review process to potentially obtain better captions by looking into semantic components with a global view, where relations can also play the role of the prompt signal for the controllable caption generation at the second decoding pass. Extensive experiments show that our method achieves the SOTA performance on MSVD, MSRVTT, and VATEX datasets for video captioning, and is capable of controlling captions to respond to different semantic contexts.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- From Skeleton to Flesh: Aggregated Relational Transformer Towards Controllable Video Captioning with Two-Step Decoding
- Date Crossref
- 30/06/2025
- Éditeur
- ACM
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
China University of Petroleum pays non établi dans la noticeUniversité ou école supérieure
-
Beijing Institute of Technology pays non établi dans la noticeUniversité ou école supérieure
-
Bytedance pays non établi dans la noticeInstitution
China University of Petroleum, Beijing Institute of Technology et Bytedance.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.