HumanCLAW: best frontier VLM completes only 16.8% of embodied interaction tasks
jiqizhixin · x · 2026-09-13
Meta, NTU, UW, Brown, and Northwestern present HumanCLAW, testing whether vision-language models can actually act through a body rather than just describe what they see.
Key points:
- Separates high-level action selection from low-level motor control while preserving continuous motion, collision, contact, and gravity
- A frozen off-the-shelf VLM makes decisions from a first-person view, choosing each step from parameterized atomic skills (move forward, turn, sit down)
- The framework studies "what to do", leaving joint-level control to a reliable low-level controller
Result: nine frontier VLMs tested under the same protocol — the best completes only 16.8% of full interactions, a stark finding that understanding the world is not the same as acting in it. Fully open-sourced.
More from Embodied
- Digital fly with its own connectome brain perceives a real room via AR glasses — stspanho · 2026-09-13
- One image plus one prompt: RL robot demo built in half an hour — heatherknight · 2026-09-13
- Timelines Flooded With GPT-6 Astra Robot Demos as Physical AI Hypes Up — CyberRobooo · 2026-09-13
- DLSS 5 gives Madden its first graphics boost in a decade, fans joke devs can coast another 10 years — chrisfirst · 2026-09-13
- A humanoid robot casually pushes the grease bot out of its way — MatthewChang · 2026-09-13
- Unitree G1 teardown: how this humanoid robot's actuators move it — ericjang11 · 2026-09-13