HumanCLAW: Top VLMs Lack Embodied Self-Awareness, Score Only 16.8%
metaresearch · hf · 2026-07-30
Evaluating whether a Vision-Language Model (VLM) can act through a physical body is challenging, as task failures often blur the line between poor decision-making and low-level execution errors like losing balance.
This work introduces HumanCLAW, an evaluation framework that decouples action decision-making from physical execution. The VLM issues atomic skill commands translated into full-body motions with real physical consequences (gravity and collisions), factoring out motor errors to purely measure the model's action intelligence.
Researchers built the HumanCLAW-Bench featuring 1,218 long-horizon episodes across 41 indoor scenes. Testing nine state-of-the-art VLMs revealed that none could solve the benchmark, with the best model achieving only a 16.8% success rate.
The study concludes that target recognition is not the bottleneck. Instead, current VLMs critically lack embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or if it has hit an obstacle.
More from Embodied
- Waymo's new UI in Ojai gets praised as 10x improvement, massive step up — brianwilt · 2026-07-30
- Ropedia Raises $30M to Build Real-World Data Infrastructure for Physical AI — liuziwei7 · 2026-07-30
- Langostino: Open-Source Autonomous Drone with ROS2 and RL — tom_doerr · 2026-07-30
- REGRIND: Humanoid Robots Learn Tool Use from a Single Human Demonstration — HaozhiQ · 2026-07-30
- TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on RTX 4090 — H-EmbodVis · 2026-07-30
- Unmasking a Common Trick in Robot Demos: Pre-loading Objects in Grippers — chris_j_paxton · 2026-07-30