Aller au contenu principal
2025 article

HeadArtist-VL: Vision / Language Guided 3D Head Generation with Self Score Distillation

0Citations signalées, ce qui n’est pas une note de qualité
3Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : hk, cn. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

We present HeadArtist-VL, a 3D head generation method that suits either vision or language input. With a landmark-guided ControlNet serving as a generative prior, we come up with an efficient pipeline that optimizes a parameterized 3D head model under the supervision of the prior distillation itself. We name such a process self-score distillation (SSD). In detail, given a sampled camera pose, we first render an image and its corresponding landmarks from the head model, and add some particular level of noise onto the image. When the input is a language prompt, we fed the noisy image, landmarks, and the language prompt into a frozen ControlNet twice for noise prediction. We conduct two predictions via the same ControlNet structure but with different classifier-free guidance (CFG) weights. The difference between these two predicted results directs how the rendered image can better match the language instructions. When the input is a reference image, we follow the aforementioned pipeline but with two modifications. First, we use an image encoder to obtain the image identity embedding, which is then sent to the ControlNet. Second, we use a novel-view diffusion model to synthesize the same reference image under the sampled camera pose to guide the self-score distillation process. In the experiments, our HeadArtist-VL produces high-quality 3D head sculptures with rich geometry and photo-realistic appearance, which significantly outperforms state-of-the-art methods. We also show that our method supports editing operations on the generated heads, including both geometry deformation and appearance change. 3D Head Generation and Editing, Vision / Language Guidance, Self Score Distillation.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
HeadArtist-VL: Vision / Language Guided 3D Head Generation with Self Score Distillation
Date Crossref
01/01/2025
Éditeur
Institute of Electrical and Electronics Engineers (IEEE)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Hong Kong University of Science and Technology HKUST pays non établi dans la notice
    Université ou école supérieure
  • Ant Group (China) pays non établi dans la notice
    Entreprise
  • City University of Hong Kong The Department of Computer Science pays non établi dans la notice
    Université ou école supérieure

HKUST — Hong Kong University of Science and Technology, Ant Group (China) et The Department of Computer Science — City University of Hong Kong.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Multimodal Machine Learning ApplicationsSocial Robot Interaction and HRIAI in Service Interactions

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.