Aller au contenu principal
2024 article

Still-Moving: Customized Video Generation without Customized Video Data

22Citations signalées, ce qui n’est pas une note de qualité
6Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : il, us. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Customizing text-to-image (T2I) models has seen tremendous progress recently, particularly in areas such as personalization, stylization, and conditional generation. However, expanding this progress to video generation is still in its infancy, primarily due to the lack of customized video data. In this work, we introduce Still-Moving, a novel generic framework for customizing a text-to-video (T2V) model, without requiring any customized video data. The framework applies to the prominent T2V design where the video model is built over a T2I model (e.g., via inflation). We assume access to a customized version of the T2I model, trained only on still image data (e.g., using DreamBooth). Naively plugging in the weights of the customized T2I model into the T2V model often leads to significant artifacts or insufficient adherence to the customization data. To overcome this issue, we train lightweight Spatial Adapters that adjust the features produced by the injected T2I layers. Importantly, our adapters are trained on "frozen videos" (i.e., repeated images), constructed from image samples generated by the customized T2I model. This training is facilitated by a novel Motion Adapter module, which allows us to train on such static videos while preserving the motion prior of the video model. At test time, we remove the Motion Adapter modules and leave in only the trained Spatial Adapters. This restores the motion prior of the T2V model while adhering to the spatial prior of the customized T2I model. We demonstrate the effectiveness of our approach on diverse tasks including personalized, stylized, and conditional generation. In all evaluated scenarios, our method seamlessly integrates the spatial prior of the customized T2I model with a motion prior supplied by the T2V model.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Still-Moving: Customized Video Generation without Customized Video Data
Date Crossref
19/11/2024
Éditeur
Association for Computing Machinery (ACM)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Tel Aviv University pays non établi dans la notice
    Université ou école supérieure
  • Google (Israel) pays non établi dans la notice
    Entreprise
  • Boston University pays non établi dans la notice
    Université ou école supérieure
  • Google (United States) pays non établi dans la notice
    Entreprise
  • Weizmann Institute of Science pays non établi dans la notice
    Université ou école supérieure
  • Technion – Israel Institute of Technology pays non établi dans la notice
    Université ou école supérieure
  • Google Research pays non établi dans la notice
    Institution
  • Technion- Israel Institute of Technology pays non établi dans la notice
    Structure de recherche

Tel Aviv University, Google (Israel) et Boston University, avec 5 autres affiliations.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Cinema and Media StudiesVideo Analysis and SummarizationGenerative Adversarial Networks and Image Synthesis

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.