Meta's HumanCLAW benchmark: top VLMs fail badly at acting through a body across 1,218 tasks
wzenus · x · 2026-09-08
Researchers from Meta, NTU, UW and others released HumanCLAW, asking whether VLMs can act through a body — e.g., walk over to a couch and sit down. The work introduces "Action Intelligence," decoupled from motor control so it can be measured cleanly: a frozen VLM picks one parametric whole-body skill every 0.5s from an egocentric view, and a pretrained motion generator turns it into continuous motion, so failure = decision failure.
The benchmark spans 1,218 long-horizon tasks across 41 indoor scenes, and today's best models are surprisingly bad. Key failure modes: inefficient exploration (recognizing targets but never bringing them into view) and reaching-but-not-arriving (walls, feet catching on furniture, knocking objects aside).
Paper, code and leaderboard are open.
More from Embodied
- Arm's CSS for Mobile 2 Skips the NPU, Bets On-Device AI on CPU and GPU Engines — ryanshrout · 2026-09-08
- Non-engineer builds a modern HP Jornada palmtop with GPT-6 Astra doing the engineering reasoning — Wuliyasi · 2026-09-08
- Weekend Hack Pays Off: DIY Robot Achieves Autonomous Obstacle Avoidance — _Stocko_ · 2026-09-08
- XPENG's IRON becomes first humanoid to walk off its own production line, targeting 1,000 units/month — CyberRobooo · 2026-09-08
- GPT-6 Astra scores 95% on robot control task — but critics say demos are gamed — GaryMarcus · 2026-09-08
- First Cybercab ride: 'exceeded my expectations' with surprisingly roomy interior — jevon · 2026-09-08