LIBERO-MAX: mid-task changes cut success rates for all 14 robot policies by 11-25.7 points
DJiafei · x · 2026-10-08
Researchers released LIBERO-MAX, a dynamic robustness benchmark testing whether robot policies can still finish tasks when the world changes mid-execution — e.g., the cup moves after the arm commits.
- 8,000 paired cases: every Dynamic rollout is matched to a Base control sharing task, seed, and executed prefix, isolating the effect of one injected change event.
- 8 event types, 1,000 cases each: target/receptacle relocation, camera shift, sensor corruption, illumination switch, visual theme change, distractor burst, obstacle insertion.
- Unlike static benchmarks (LIBERO-Plus/PRO) that test shifted resets, this measures adaptation after the policy has committed; it does not infer internal detection or deliberate replanning.
Across all 14 tested robot policies, success drops by 11.0–25.7 percentage points. Code, dataset, and paper are public.
More from Embodied
- Undergrad attends first IROS, gets hands-on with new humanoid robots — heatherknight · 2026-10-08
- Desktop robots MOSS and Vita star in an 80-minute robotics seminar at Thammasat University — StewartalsopIII · 2026-10-08
- Nvidia bets billions on physical AI, expanding Halos safety system to robots — Ars Technica AI · 2026-10-08
- Singapore debuts real-life Rock 'Em Sock 'Em robot boxing match — DJiafei · 2026-10-08
- KAIST's FastOPD cuts VLA robot inference latency by 78% via on-policy distillation — kaist-ai · 2026-10-08
- Meta's 8B RoboJEPA shows robot world models have scaling laws, trained on 15k hours of video — ylecun · 2026-10-08