Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding
TL;DR AI
2 min readKey summary
The paper presents Think, Act, Build (TAB), an agentic framework that uses 2D VLMs plus multi-view geometry for 3D visual grounding.
TAB operates on raw RGB-D streams and reformulates 3D-VG as a 2D-to-3D reconstruction task driven by a VLM agent.
The authors introduce Semantic-Anchored Geometric Expansion to propagate target locations across frames and aggregate multi-view features into 3D coordinates.
Evaluations on ScanRefer and Nr3D show TAB outperforms prior zero-shot methods and exceeds some supervised baselines.
