Leveraging Latent Visual Reasoning in Silence

TL;DR AI
2 min readKey summary
Researchers found that latent visual tokens can be removed at inference with little loss on several benchmarks, challenging the idea that they must always stay visible.
The tokens still help multimodal models learn better visual grounding during training, especially for visual and spatial reasoning.
To amplify that benefit, the team introduced an attention-based reinforcement learning reward.
The approach improves visual-text reasoning while preserving strong pure-text reasoning when visual tokens are not needed.
