HumanCLAW: Can Vision-Language Models Act Through a Body?
TL;DR AI
2 min readKey summary
Researchers introduced HumanCLAW, a benchmark that turns vision-language model commands into short full-body motion chunks to test embodied control.
They also built HumanCLAW-Bench, with 1,218 egocentric indoor episodes across 41 scenes, to measure real-world physical decision-making.
Testing nine leading VLMs, none solved the benchmark; the best model reached only 16.8% success.
The main failure was not target recognition, but weak embodied self-awareness and poor tracking of the body over long-horizon tasks.
