Accès ouvert
2026
preprint
OpenAlex
Yiheng Lin, Siyu Jiao, Xiaohan Lan, Zhou, Wei, 1968 Jan. 17- et autres
Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editing. However, scene text editing remains challenging, as it requires models to precisely modify textual content while preserving visual realism and non-target regions. Current open-source …
Accès ouvert
2026
preprint
OpenAlex
Yiheng Lin, Siyu Jiao, Xiaohan Lan, Wei Zhou et autres
Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editing. However, scene text editing remains challenging, as it requires models to precisely modify textual content while preserving visual realism and non-target regions. Current open-source …
2025
conference-paper
OpenAlex
Yuji Wang, Moran Li, Xiaobin Hu, Ran Yi et autres
Identity-preserving text-to-video (IPT2V) generation, which aims to create high-fidelity videos with consistent human identity, has become crucial for downstream applications. However, current end-to-end frameworks suffer a critical spatial-temporal trade-off: optimizing for spatially coherent layouts of key elements ( e.g., character identity preservation) …
cn
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Yuji Wang, Moran Li, Xiaobin Hu, Ran Yi et autres
Current video generation models struggle with identity preservation under large face poses, primarily facing two challenges: the difficulty in exploring an effective mechanism to integrate identity features into DiT architectures, and the lack of targeted coverage of large face poses in existing …
2024
conference-paper
OpenAlex
Ke Fan, Junshu Tang, Weijian Cao, Ran Yi et autres
cn
(code pays fourni par la source)
Accès ouvert
2023
article
OpenAlex
Yuan Cui, Moran Li, Yuan Gao, Changxin Gao et autres
Most existing methods for RGB hand pose estimation use root-relative 3D coordinates for supervision. However, such supervision neglects the distance between the camera and the object (i.e., the hand). The camera distance is especially important under a perspective camera, which controls the …
cn
(code pays fourni par la source)
2022
article
OpenAlex
Moran Li, Haibin Huang, Yi Zheng, Mengtian Li et autres
Abstract In this work, we present a new method for 3D face reconstruction from sparse‐view RGB images. Unlike previous methods which are built upon 3D morphable models (3DMMs) with limited details, we leverage an implicit representation to encode rich geometric features. Our …
cn
(code pays fourni par la source)
2021
article
OpenAlex
Moran Li, Jialong Wang, Nong Sang
In this article, we propose a novel compressed latent distribution representation for 3D hand pose estimation from monocular RGB images to alleviate the channel correspondence problem. The channel correspondence problem occurs when the 2D and depth coordinates are estimated from independent feature …
cn
(code pays fourni par la source)