HeadArtist-VL: Vision / Language Guided 3D Head Generation with Self Score Distillation
Rattachement africain : hk, cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
We present HeadArtist-VL, a 3D head generation method that suits either vision or language input. With a landmark-guided ControlNet serving as a generative prior, we come up with an efficient pipeline that optimizes a parameterized 3D head model under the supervision of the prior distillation itself. We name such a process self-score distillation (SSD). In detail, given a sampled camera pose, we first render an image and its corresponding landmarks from the head model, and add some particular level of noise onto the image. When the input is a language prompt, we fed the noisy image, landmarks, and the language prompt into a frozen ControlNet twice for noise prediction. We conduct two predictions via the same ControlNet structure but with different classifier-free guidance (CFG) weights. The difference between these two predicted results directs how the rendered image can better match the language instructions. When the input is a reference image, we follow the aforementioned pipeline but with two modifications. First, we use an image encoder to obtain the image identity embedding, which is then sent to the ControlNet. Second, we use a novel-view diffusion model to synthesize the same reference image under the sampled camera pose to guide the self-score distillation process. In the experiments, our HeadArtist-VL produces high-quality 3D head sculptures with rich geometry and photo-realistic appearance, which significantly outperforms state-of-the-art methods. We also show that our method supports editing operations on the generated heads, including both geometry deformation and appearance change. 3D Head Generation and Editing, Vision / Language Guidance, Self Score Distillation.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- HeadArtist-VL: Vision / Language Guided 3D Head Generation with Self Score Distillation
- Date Crossref
- 01/01/2025
- Éditeur
- Institute of Electrical and Electronics Engineers (IEEE)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Hong Kong University of Science and Technology HKUST pays non établi dans la noticeUniversité ou école supérieure
-
Ant Group (China) pays non établi dans la noticeEntreprise
-
City University of Hong Kong The Department of Computer Science pays non établi dans la noticeUniversité ou école supérieure
HKUST — Hong Kong University of Science and Technology, Ant Group (China) et The Department of Computer Science — City University of Hong Kong.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.