Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models

TL;DR AI
2 min readKey summary
Researchers introduced a new spatial reasoning pipeline for vision-language models using a dynamic cognitive map as persistent scene memory.
The system also adds Spatial Assertion Codes to verify intermediate reasoning steps, making the process more checkable and robust.
Trained with supervised and reinforcement finetuning, it achieved state-of-the-art results on MindCube with 80.5% overall accuracy.
The biggest gain came on the Rotation subset, suggesting the approach could improve agentic perception systems in real-world settings.
