Meta's HumanCLAW benchmark: top VLMs fail badly at acting through a body across 1,218 tasks

wzenus · x · 2026-09-08

Researchers from Meta, NTU, UW and others released HumanCLAW, asking whether VLMs can act through a body — e.g., walk over to a couch and sit down. The work introduces "Action Intelligence," decoupled from motor control so it can be measured cleanly: a frozen VLM picks one parametric whole-body skill every 0.5s from an egocentric view, and a pretrained motion generator turns it into continuous motion, so failure = decision failure.

The benchmark spans 1,218 long-horizon tasks across 41 indoor scenes, and today's best models are surprisingly bad. Key failure modes: inefficient exploration (recognizing targets but never bringing them into view) and reaching-but-not-arriving (walls, feet catching on furniture, knocking objects aside).

Paper, code and leaderboard are open.

Original post →

More from Embodied

Embodied channel →