Aller au contenu principal
2026 article

Training-Free Controllable Text-Guided Video Editing

1Citations signalées, ce qui n’est pas une note de qualité
4Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Decomposition-based text-guided video editing paradigm aims to utilize the layered neural atlas model to decompose the input video into foreground and background parts and edit the video in a divide-and-conquer manner, which is meaningful and improves the controllability of editing. However, they may suffer from some limitations: 1) High computational cost of per-video training (i.e, 7∼8 hours for training a single atlas model). 2) Foreground object deformation is restricted by the foreground opacity value. 3) Restricted flexibility in manipulating multiple objects. In this paper, we propose TraFrCo, aTraining-Free Controllable Text-guided Video Editingframework to mitigate these challenges. Instead of training complex atlas models, our method leverages pre-trained segmentation to rapidly decompose videos into foreground and background parts. This allows users to perform independent edits on foreground objects using existing video diffusion editing models without affecting the environment. To ensure visual consistency, we introduce a training-free mechanism that effectively propagates information across frames to fill missing background regions caused by the segmentation-derived foreground masks and reconstructs the scene behind moving objects. Finally, the edited components are seamlessly composited by re-predicting the new foreground masks. In contrast to prior works, TraFrCo enables efficient, fine-grained manipulation of video content without the burden of training. Experimental results verify that our TraFrCo consistently reduces the costs of decomposing video and achieves superior text-guided video editing performance. Codes and video demos will be released at https://github.com/mdswyz/TraFrCo.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Training-Free Controllable Text-Guided Video Editing
Date Crossref
01/05/2026
Éditeur
Institute of Electrical and Electronics Engineers (IEEE)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Generative Adversarial Networks and Image SynthesisVideo Analysis and SummarizationMultimodal Machine Learning Applications

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.