SceneActBench: Can Agents Act on the 3D Scenes They See?
TL;DR AI
2 min readKey summary
SceneActBench is a new benchmark for evaluating how vision-language agents act in complete 3D scenes.
It tests five visually conditioned 3D tasks in a fixed loop using images or video frames, and sometimes 3D assets.
The benchmark scores final actions with hidden ground truth and geometric metrics across 520 cases from 210 source instances.
Results show that eleven proprietary VLMs still perform unevenly across tasks, highlighting a gap in real scene-action evaluation.
