Aller au contenu principal
Accès ouvert déclaré 2026 article

Beyond VQA Accuracy: A Cross-Regime Diagnostic Evaluation of Symbol Consistency in Multimodal Large Language Models

0Citations signalées — pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Résumé fourni par la source

Multimodal large language models (MLLMs) have achieved strong performance in general visual question answering, yet their visual faithfulness in handling discrete symbolic information remains insufficiently understood. Symbols such as numbers, time expressions, identifiers, license plates, and alphanumeric strings impose low semantic redundancy and strict character-level constraints, making overall VQA accuracy or isolated OCR-style evaluation inadequate for diagnosing model reliability. To address this gap, this paper proposes a cross-regime diagnostic framework for evaluating symbol consistency in MLLMs. Under a unified protocol, we evaluate five representative models, including BLIP-2, InstructBLIP, LLaVA, InternVL, and Qwen, across VQA, TextVQA, a self-constructed Symbol Subset, Regime 2a with uncontrolled generative symbol rendering, and Regime 2b with controlled clear-symbol grounding. We further introduce a 2a $\rightarrow 2$ b paired recovery analysis to distinguish rendering-sensitive errors caused by upstream symbol degradation from persistent errors that remain under clear visual evidence. Regime 2a is interpreted as an uncontrolled stress probe rather than a clean OCR benchmark, and its accuracy reflects both upstream rendering quality and downstream model reading. Experimental results show that general VQA accuracy can substantially overestimate symbol-centered reliability, especially for weaker models. Symbol consistency failures are not merely OCR recognition errors, but arise from the interaction of target-region binding, character-faithful transcription, answer completeness, and language-prior normalization. Although clear-symbol conditions improve stronger models, persistent failures remain in character-level grounding, target binding, and task following. This study separates symbol consistency from general VQA evaluation and reveals its multi-stage failure mechanisms, offering a reusable diagnostic perspective for reliability assessment and symbol-capability improvement in MLLMs.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Beyond VQA Accuracy: A Cross-Regime Diagnostic Evaluation of Symbol Consistency in Multimodal Large Language Models
Date Crossref
01/01/2026
Éditeur
Institute of Electrical and Electronics Engineers (IEEE)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Topic ModelingNatural Language Processing TechniquesText Readability and Simplification

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, ROR et la Banque mondiale, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune donnée externe enregistrée en base. Sources et limites.