HumanCLAW: Top VLMs Lack Embodied Self-Awareness, Score Only 16.8%

metaresearch · hf · 2026-07-30

Evaluating whether a Vision-Language Model (VLM) can act through a physical body is challenging, as task failures often blur the line between poor decision-making and low-level execution errors like losing balance.

This work introduces HumanCLAW, an evaluation framework that decouples action decision-making from physical execution. The VLM issues atomic skill commands translated into full-body motions with real physical consequences (gravity and collisions), factoring out motor errors to purely measure the model's action intelligence.

Researchers built the HumanCLAW-Bench featuring 1,218 long-horizon episodes across 41 indoor scenes. Testing nine state-of-the-art VLMs revealed that none could solve the benchmark, with the best model achieving only a 16.8% success rate.

The study concludes that target recognition is not the bottleneck. Instead, current VLMs critically lack embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or if it has hit an obstacle.

Original post →

More from Embodied

Embodied channel →