RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models
TL;DR AI
2 min readKey summary
Researchers introduced RoboSemanticBench, an embodied benchmark that asks robots to answer multiple-choice math and general-knowledge questions by grasping the block with the correct answer.
When tested on representative vision-language-action models, many robots could physically grasp objects but selected the semantically correct block at near-random or even worse-than-random rates once grasp success was separated out.
The results highlight a major gap between pretrained language knowledge and action prediction in robot policies.
RoboSemanticBench exposes a weakness in semantic grounding: robots may execute grasps well, yet still fail to map instructions to the right target.
