Switch language한국어
Back to the list

Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels

TL;DR AI

Key summary

2 min read
  1. Researchers proposed a language-based interface for visual document QA that predicts verbatim quotes as evidence instead of coordinate boxes.

  2. A parser and retriever then map those quotes to page regions, enabling training without region-level evidence labels.

  3. On CiteVQA, evidence recall jumped from at most 8% to about 47%, while attribution hallucination was cut roughly in half.

  4. With GRPO training, strict attributed accuracy improved from 22.4% to 33.8%.

Read the original