Seeing Before Answering: Training-Free Visual Layer Profiling for Vision-Language Models
Ruchen Liu, Yi Yang, Yiming Xu, Michael Ying Yang et autres
LLaVA-style Vision-Language Models (VLMs) pass visual tokens from a fixed late layer of the vision backbone, typically the penultimate one, to the language model. We first show that this hidden convention is fragile: across 2 VLMs and 7 image and video benchmarks, …