RobotWorld Benchmark Tests Multimodal Agents on 84 Physical Robot Tasks
Zhiqin Yang · hf · 2026-10-08
Researchers introduce RobotWorld, a simulation testbed for "robot use": converting instructions and observations into physical task execution through robot interfaces.
- 84 tasks spanning manipulation, mobile manipulation, locomotion, driving, and aerial control, with explicit interaction budgets and executable success checks.
- Uneven transfer: current agents can build sophisticated perception and control workflows (image segmentation, camera calibration, spatial estimation, dynamics-based computation) but fail to compose them reliably — losing task-relevant object states, failing to correct ineffective actions, recovering too late, or mistaking unfinished tasks for completion.
- Model differences: Astra succeeds more on spatial and constrained-contact goals, while Opus 5.5 does better on continuous-balance and timed-interaction goals.
The benchmark provides concrete targets for training more reliable physical-world agents.
More from Embodied
- HKUST-GZ's UniWAM Unifies Physical Reasoning, World Generation and Action Prediction, Finds Co-Training Scaling Law — HKUSTGZ · 2026-10-08
- PKU's ViGAR Hierarchical World-Action Model Boosts Robot Manipulation Success by 12.86 Points — DAGroup-PKU · 2026-10-08
- Open-source agentic workbench: a harness that hears, speaks and sees to help assemble electronics — andreisavu · 2026-10-08
- RoboQuest Benchmark: Best Multimodal Agent Succeeds in Only 23% of Embodied Exploration Tasks — declare-lab · 2026-10-08
- NVIDIA's Long-WAM scales world-action model context, hitting 95% on dynamic cup stacking — nvidia · 2026-10-08
- Donut Robotics tests 170cm humanoid in Japanese eldercare across 100+ facilities — CyberRobooo · 2026-10-08