iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning
TL;DR AI
2 min readKey summary
Researchers introduced iVGR, a reinforcement learning framework that transfers visual grounding into textual reasoning for multimodal LLMs.
They found that requiring explicit object boxes at inference can hurt performance, so they trained a dual-stream system with a consistency reward.
The method aligns Chain-of-Thought reasoning with a visually grounded stream, improving fine-grained perception on benchmarks.
iVGR boosts multimodal understanding while still supporting tool-assisted workflows without needing explicit grounding at inference.
