VA-Bench: Best MLLM Scores Only 53.9% Task Success on Embodied Spatial Intelligence

dalian-university-of-technology · hf · 2026-09-18

Dalian University of Technology introduces VA-Bench, a benchmark evaluating the full observe-reason-act-revise loop of MLLMs under incomplete observation: models must actively select camera viewpoints, issue metric Cartesian commands, and revise from execution feedback, with no privileged object poses, oracle trajectories, or learned action heads.

The benchmark spans 14 base task families (11 single-arm, three dual-arm), seven held-out geometry/layout variants, and a long-horizon five-object composition track, with 12 primary model conditions evaluated over three runs on 20 physically verified seeds each.

Key findings:

The takeaway: general-purpose MLLMs remain far from turning visual demonstrations and actively acquired evidence into successful embodied action.

Original post →

More from Embodied

Embodied channel →