How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
TL;DR AI
2 min readKey summary
Researchers propose View Dropout, which hides parts of one view from the answer path while keeping them visible for visual thinking tokens.
They compare top-down, panoramic, and point-matching visual reasoning, and find panoramic thinking is the easiest to learn and most useful for reasoning.
View Dropout plus panoramic visual thinking is the only setup that consistently improves cross-view spatial reasoning.
The approach delivers the best results on five out-of-domain benchmarks, showing stronger generalization beyond training conditions.
