TVIR: Building Deep Research Agents Towards Text-Visual Interleaved Report Generation
TL;DR AI
2 min readKey summary
Researchers introduced TVIR, a multimodal benchmark and agent framework for generating reports that integrate text and visuals with stronger factual and visual verification.
TVIR-Bench includes 100 expert-curated tasks that require images to support specific analytical goals, filling a key gap in deep research evaluation.
TVIR-Agent uses a hierarchical multi-agent pipeline for outlining, image retrieval, chart generation with traceable sources, and sequential report writing.
The authors also proposed dual-path evaluation that assesses both textual quality and visual alignment, and tested it across nine deep research systems.
