Aller au contenu principal
Accès ouvert déclaré2026preprint

Human-like meaning maps from single-prompt VLM ratings of local scene meaning

0Citations signalées
5Institutions associées
2Pays d’affiliation

Résumé fourni par la source

Abstract Visual attention is shaped both by low-level perceptual features and by the semantic properties of scene regions. However, compared to the salience of low-level image features, the characterization of image properties that lead to scene understanding has been more difficult. The semantic properties of scene regions have been operationalized via meaning maps, proposed by Henderson and Hayes (2017), by having human observers rate the informativeness of scene patches. However, this rating procedure is labor-intensive and limits the method’s versatility. An automated alternative, the DeepMeaning model, reduces rating costs but still requires human training data from a similar class of images, limiting generalizability. Here, we tested whether instruction-following vision-language models (VLMs) can generate human-like ratings in a zero-shot manner: using the original human-rating instructions as a single prompt without any task-specific training. We obtained meaningfulness ratings from both open-weight and proprietary VLMs for scene patches across three meaning-map datasets, and compared the resulting zero-shot AI meaning maps (AIMMs) to human meaning maps and maps obtained from the DeepMeaning model. Zero-shot AIMMs closely approximate human meaning maps, and maps from the task-specific DeepMeaning model, showing that VLMs can replicate human semantic judgments without any training. Zero-shot AIMMs demonstrate utility and validity, matching human meaning maps in their ability to predict human fixations and reproducing the same pattern of relative performance against other gaze-prediction models. We provide open-source code that allows researchers to use this zero-shot approach to quickly and flexibly explore meaning maps across spatial scales, prompts, and scene categories.

Institutions

Sujets associés

Visual Attention and Saliency DetectionMultimodal Machine Learning ApplicationsGaze Tracking and Assistive Technology

BNTIC News n’est pas le producteur de ces données. Métadonnées interrogées à la demande auprès de OpenAlex (CC0). Sources et limites.