Accès ouvert
2026
preprint
OpenAlex
Gengyuan Liu, Nanzhou Wang, Chang Liu, Qinwen Wu et autres
Vision-language models exhibit remarkable zero-shot capabilities but suffer significant performance degradation under distribution shifts. While test-time adaptation (TTA) via Low-Rank Adaptation offers a parameter-efficient solution, we identify a fundamental bottleneck in current methods: the reliance on static rank configurations. Because visual inputs …
cn
(code pays fourni par la source)
Accès ouvert
2025
conference-paper
OpenAlex
Weixian Lei, Jiacong Wang, Haochen Wang, Xiangtai Li et autres
This paper introduces SAIL, a single transformer unified multimodal large language model (MLLM) that integrates raw pixel encoding and language decoding within a singular architecture. Unlike existing modular MLLMs, which rely on a pre-trained vision transformer (ViT), SAIL eliminates the need for …
pl
(code pays fourni par la source)
Accès ouvert
2025
conference-paper
OpenAlex
Yuanxin Liu, Rui Zhu, Shuhuai Ren, Jiacong Wang et autres
With the rapid growth of video generative models (VGMs), it is essential to develop reliable and comprehensive automatic metrics for AI-generated videos (AIGVs). Existing methods either use off-the-shelf models optimized for other tasks or rely on human assessment data to train specialized …
cn, gb
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Hongyuan Dong, Jiawen Li, Bohong Wu, Jiacong Wang et autres
Image captioning has long been regarded as a fundamental task in visual understanding. Recently, however, few large vision-language model (LVLM) research discusses model's image captioning performance because of the outdated short-caption benchmarks and unreliable evaluation metrics. In this work, we propose to …
Accès ouvert
2024
preprint
OpenAlex
Xin Xiao, Bohong Wu, Jiacong Wang, Chunyuan Li et autres
Existing image-text modality alignment in Vision Language Models (VLMs) treats each text token equally in an autoregressive manner. Despite being simple and effective, this method results in sub-optimal cross-modal alignment by over-emphasizing the text tokens that are less correlated with or even …
Accès ouvert
2024
preprint
OpenAlex
Yuan Zhang, Fei Xiao, Tao Huang, Chun-Kai Fan et autres
Large vision-language models (LVLMs) have recently achieved rapid progress, exhibiting great perception and reasoning abilities concerning visual information. However, when faced with prompts in different sizes of solution spaces, LVLMs fail to always give consistent answers regarding the same knowledge point. This …
2024
conference-paper
OpenAlex
Fei Xiao, Tao Huang, Chun-Kai Fan, Hongyuan Dong et autres
2024
conference-paper
OpenAlex
Xin Xiao, Bohong Wu, Jiacong Wang, Xun Zhou et autres
2023
article
OpenAlex
Jiacong Wang, Xiaolan Ding, Jun Xiao
cn
(code pays fourni par la source)
2022
conference-paper
OpenAlex
Jiacong Wang, Cheng Tang, Jianping Li
Rapid and quantitative analysis of phytoplankton cells in natural seawater is of great need for marine ecological science research and harmful algae bloom monitoring applications. In this paper, we propose a YOLOX network-based object detection algorithm exclusively for high-throughput real-time analysis of …
cn
(code pays fourni par la source)