VA-Bench: Best MLLM Scores Only 53.9% Task Success on Embodied Spatial Intelligence
dalian-university-of-technology · hf · 2026-09-18
Dalian University of Technology introduces VA-Bench, a benchmark evaluating the full observe-reason-act-revise loop of MLLMs under incomplete observation: models must actively select camera viewpoints, issue metric Cartesian commands, and revise from execution feedback, with no privileged object poses, oracle trajectories, or learned action heads.
The benchmark spans 14 base task families (11 single-arm, three dual-arm), seven held-out geometry/layout variants, and a long-horizon five-object composition track, with 12 primary model conditions evaluated over three runs on 20 physically verified seeds each.
Key findings:
- The best model hits 100.0% on target localization and 78.9% on spatial relations in the annotated run, but its three-run macro-average task success is only 53.93±3.17%
- Active camera control beats passive multi-view observation: one matched comparison rises from 27.86% to 57.50% success
- Held-out geometric transfer can cut task success by over 30 percentage points
- No model completes a strict long-horizon episode despite substantial partial progress
The takeaway: general-purpose MLLMs remain far from turning visual demonstrations and actively acquired evidence into successful embodied action.
More from Embodied
- DARPA surgical AI competition kicks off, veterans liken it to the 2004 self-driving challenge — Laparoscopes · 2026-09-18
- visloc-rs: Pure Rust library for visual & visual-inertial SLAM, SfM and localization — rsasaki0109 · 2026-09-18
- Scoble Teases Camera-Free, Display-Free All-Day AI Glasses Launching Next Week — Scobleizer · 2026-09-18
- Waymo Announces Its Second Asia City for Robotaxi Expansion — reed · 2026-09-18
- SceneAgent: agentic pipeline turns 3D captures into physics-ready scenes for robot training — hankyang94 · 2026-09-18
- DoorDash's Dot robot is autonomously delivering real orders daily — ycombinator · 2026-09-18