Beyond VQA Accuracy: A Cross-Regime Diagnostic Evaluation of Symbol Consistency in Multimodal Large Language Models
Résumé fourni par la source
Multimodal large language models (MLLMs) have achieved strong performance in general visual question answering, yet their visual faithfulness in handling discrete symbolic information remains insufficiently understood. Symbols such as numbers, time expressions, identifiers, license plates, and alphanumeric strings impose low semantic redundancy and strict character-level constraints, making overall VQA accuracy or isolated OCR-style evaluation inadequate for diagnosing model reliability. To address this gap, this paper proposes a cross-regime diagnostic framework for evaluating symbol consistency in MLLMs. Under a unified protocol, we evaluate five representative models, including BLIP-2, InstructBLIP, LLaVA, InternVL, and Qwen, across VQA, TextVQA, a self-constructed Symbol Subset, Regime 2a with uncontrolled generative symbol rendering, and Regime 2b with controlled clear-symbol grounding. We further introduce a 2a $\rightarrow 2$ b paired recovery analysis to distinguish rendering-sensitive errors caused by upstream symbol degradation from persistent errors that remain under clear visual evidence. Regime 2a is interpreted as an uncontrolled stress probe rather than a clean OCR benchmark, and its accuracy reflects both upstream rendering quality and downstream model reading. Experimental results show that general VQA accuracy can substantially overestimate symbol-centered reliability, especially for weaker models. Symbol consistency failures are not merely OCR recognition errors, but arise from the interaction of target-region binding, character-faithful transcription, answer completeness, and language-prior normalization. Although clear-symbol conditions improve stronger models, persistent failures remain in character-level grounding, target binding, and task following. This study separates symbol consistency from general VQA evaluation and reveals its multi-stage failure mechanisms, offering a reusable diagnostic perspective for reliability assessment and symbol-capability improvement in MLLMs.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Beyond VQA Accuracy: A Cross-Regime Diagnostic Evaluation of Symbol Consistency in Multimodal Large Language Models
- Date Crossref
- 01/01/2026
- Éditeur
- Institute of Electrical and Electronics Engineers (IEEE)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.