ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
TL;DR AI
2 min readKey summary
Researchers introduced ESI-Bench, a benchmark for embodied spatial intelligence built in OmniGibson to test active evidence gathering through exploration.
The suite spans 10 task categories and 29 subcategories, focusing on how agents use perception, movement, and manipulation to uncover hidden spatial information.
Experiments show active exploration outperforms passive viewing, while random extra viewpoints can actually reduce performance.
Current MLLMs often choose actions poorly, commit too confidently to wrong beliefs, and only benefit from explicit 3D grounding when that grounding is accurate.
