Accès ouvert
2026
preprint
OpenAlex
Yiheng Lin, Siyu Jiao, Xiaohan Lan, Zhou, Wei, 1968 Jan. 17- et autres
Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editing. However, scene text editing remains challenging, as it requires models to precisely modify textual content while preserving visual realism and non-target regions. Current open-source …
Accès ouvert
2026
preprint
OpenAlex
Yiheng Lin, Siyu Jiao, Xiaohan Lan, Wei Zhou et autres
Recent advances in Multimodal Large Language Models (MLLMs) and diffusion-based generative models have substantially improved prompt-driven image editing. However, scene text editing remains challenging, as it requires models to precisely modify textual content while preserving visual realism and non-target regions. Current open-source …
Accès ouvert
2025
preprint
OpenAlex
Siyu Jiao, Yiheng Lin, Yujie Zhong, Qi She et autres
Recent progress in Multimodal Large Language Models (MLLMs) demonstrates that Chain-of-Thought (CoT) reasoning enables systematic solutions to complex understanding tasks. However, its extension to generation tasks remains nascent and limited by scenario-specific mechanisms that hinder generalization and adaptation. In this work, we …
Accès ouvert
2025
preprint
OpenAlex
Siyu Jiao, Yiheng Lin, Yujie Zhong, Qi She et autres
Recent progress in Multimodal Large Language Models (MLLMs) demonstrates that Chain-of-Thought (CoT) reasoning enables systematic solutions to complex understanding tasks. However, its extension to generation tasks remains nascent and limited by scenario-specific mechanisms that hinder generalization and adaptation. In this work, we …
Accès ouvert
2025
conference-paper
OpenAlex
Siyu Jiao, Haoye Dong, Yuyang Yin, Zequn Jie et autres
Recent works in 3D multimodal learning have made remarkable progress. However, typically 3D multimodal models are only capable of handling point clouds. Compared to the emerging 3D representation technique, 3D Gaussian Splatting (3DGS), the spatially sparse point cloud cannot depict the texture …
cn, sg, us
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Jiancheng Huang, Gengwei Zhang, Zequn Jie, Siyu Jiao et autres
Text-to-video generation has significantly enriched content creation and holds the potential to evolve into powerful world simulators. However, modeling the vast spatiotemporal space remains computationally demanding, particularly when employing Transformers, which incur quadratic complexity in sequence processing and thus limit practical applications. …
Accès ouvert
2025
preprint
OpenAlex
Hongguang Zhu, Yunchao Wei, Mengyu Wang, Siyu Jiao et autres
Diffusion models (DMs) have achieved significant progress in text-to-image generation. However, the inevitable inclusion of sensitive information during pre-training poses safety risks, such as unsafe content generation and copyright infringement. Concept erasing finetunes weights to unlearn undesirable concepts, and has emerged as …
Accès ouvert
2025
conference-paper
OpenAlex
Siyu Jiao, Gengwei Zhang, Yinlong Qian, Jiancheng Huang et autres
This work challenges the residual prediction paradigm in visual autoregressive modeling and presents FlexVAR, a new Flexible Visual AutoRegressive image generation paradigm. FlexVAR facilitates autoregressive learning with ground-truth prediction, enabling each step to independently produce plausible images. This simple, intuitive approach swiftly …
cn, us
(code pays fourni par la source)
2024
conference-paper
OpenAlex
Xinhao Zhong, Siyu Jiao, Yao Zhao, Yunchao Wei
cn
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Xinhao Zhong, Siyu Jiao, Yao Zhao, Yunchao Wei
Current Semi-Supervised Object Detection (SSOD) methods enhance detector performance by leveraging large amounts of unlabeled data, assuming that both labeled and unlabeled data share the same label space. However, in open-set scenarios, the unlabeled dataset contains both in-distribution (ID) classes and out-of-distribution …
2024
conference-paper
OpenAlex
Siyu Jiao, Hongguang Zhu, Jiannan Huang, Yao Zhao et autres
cn, us
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Siyu Jiao, Hongguang Zhu, Jiannan Huang, Yao Zhao et autres
Pre-trained vision-language models, e.g. CLIP, have been increasingly used to address the challenging Open-Vocabulary Segmentation (OVS) task, benefiting from their well-aligned vision-text embedding space. Typical solutions involve either freezing CLIP during training to unilaterally maintain its zero-shot capability, or fine-tuning CLIP vision …