Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
TL;DR AI
2 min readKey summary
Researchers proposed a language-based interface for visual document QA that predicts verbatim quotes as evidence instead of coordinate boxes.
A parser and retriever then map those quotes to page regions, enabling training without region-level evidence labels.
On CiteVQA, evidence recall jumped from at most 8% to about 47%, while attribution hallucination was cut roughly in half.
With GRPO training, strict attributed accuracy improved from 22.4% to 33.8%.
