RoboQuest Benchmark: Best Multimodal Agent Succeeds in Only 23% of Embodied Exploration Tasks
declare-lab · hf · 2026-10-08
declare-lab introduces RoboQuest, a benchmark for goal-directed embodied exploration where agents must actively gather task-relevant information through physical interaction.
- 10 mobile manipulation tasks built on three forms of uncertainty: search, manipulation-based inspection, and interactive testing.
- Sobering results: the best of five frontier multimodal agents succeeds in only 23% of episodes; a π0.5 policy fine-tuned on released full-episode demonstrations almost never succeeds.
- Failure analysis: execution skills are not the bottleneck — agents mostly stop exploring too early, committing before observing the evidence needed for completion, and rarely prevent or repair disturbances from their own exploration. Trial-and-error learning remains difficult for most models.
The benchmark, demonstrations, and fine-tuning data are released.
More from Embodied
- HKUST-GZ's UniWAM Unifies Physical Reasoning, World Generation and Action Prediction, Finds Co-Training Scaling Law — HKUSTGZ · 2026-10-08
- PKU's ViGAR Hierarchical World-Action Model Boosts Robot Manipulation Success by 12.86 Points — DAGroup-PKU · 2026-10-08
- Open-source agentic workbench: a harness that hears, speaks and sees to help assemble electronics — andreisavu · 2026-10-08
- RobotWorld Benchmark Tests Multimodal Agents on 84 Physical Robot Tasks — Zhiqin Yang · 2026-10-08
- NVIDIA's Long-WAM scales world-action model context, hitting 95% on dynamic cup stacking — nvidia · 2026-10-08
- Donut Robotics tests 170cm humanoid in Japanese eldercare across 100+ facilities — CyberRobooo · 2026-10-08