Switch language한국어
Back to the list

HumanCLAW: Can Vision-Language Models Act Through a Body?

TL;DR AI

Key summary

2 min read
  1. Researchers introduced HumanCLAW, a benchmark that turns vision-language model commands into short full-body motion chunks to test embodied control.

  2. They also built HumanCLAW-Bench, with 1,218 egocentric indoor episodes across 41 scenes, to measure real-world physical decision-making.

  3. Testing nine leading VLMs, none solved the benchmark; the best model reached only 16.8% success.

  4. The main failure was not target recognition, but weak embodied self-awareness and poor tracking of the body over long-horizon tasks.

Read the original