HumanCLAW: best frontier VLM completes only 16.8% of embodied interaction tasks

jiqizhixin · x · 2026-09-13

Meta, NTU, UW, Brown, and Northwestern present HumanCLAW, testing whether vision-language models can actually act through a body rather than just describe what they see.

Key points:

Result: nine frontier VLMs tested under the same protocol — the best completes only 16.8% of full interactions, a stark finding that understanding the world is not the same as acting in it. Fully open-sourced.

Original post →

More from Embodied

Embodied channel →