Switch language한국어
Back to the list

RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models

TL;DR AI

Key summary

2 min read
  1. Researchers introduced RoboSemanticBench, an embodied benchmark that asks robots to answer multiple-choice math and general-knowledge questions by grasping the block with the correct answer.

  2. When tested on representative vision-language-action models, many robots could physically grasp objects but selected the semantically correct block at near-random or even worse-than-random rates once grasp success was separated out.

  3. The results highlight a major gap between pretrained language knowledge and action prediction in robot policies.

  4. RoboSemanticBench exposes a weakness in semantic grounding: robots may execute grasps well, yet still fail to map instructions to the right target.

Read the original