DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
Xiang Feng, Jiawei Zhou, Zhangfeng Huang, Kewei Wang et autres
Evaluating whether Multimodal Large Language Models can produce trustworthy, verifiable reasoning over long, visually rich documents requires evaluation beyond end-to-end answer accuracy. We introduce DocScope, a benchmark that formulates long-document QA as a structured reasoning trajectory prediction problem: given a complete PDF …