LIBERO-MAX Benchmark: Mid-Task World Changes Cut Every Robot Policy's Success by 11-26 Points
机器之心 · wechat · 2026-10-11
Researchers from Stanford, CMU, MIT, UIUC, and UT Austin introduced LIBERO-MAX, a benchmark testing whether robot policies adapt when the world changes mid-task. Accepted as an oral at the NeurIPS'26 Robotics World Modeling Workshop.
It uses paired rollouts — identical until a key moment, then one continues normally (Base) while the other faces one of 8 mid-execution perturbations (Dynamic) — across 8,000 paired cases, distinguishing pre-existing failures from change-induced ones. Evaluating 14 policies (VLAs, hybrids, world-action models) over 224,000 simulations, every policy dropped 11.0–25.7 points; the best (MolmoAct2, π₀.₅) still fell to 66%. Object/container moves and camera changes hurt most, and more frequent re-observation didn't close the gap.
A targeted fix — adding an image-quality detection and restoration module before X-VLA — lifted post-change success from 40.7% to 47.7%. Code, fixed case lists, and an 800-case Lite version are open-sourced.
More from Embodied
- Nils Pihl to keynote Humanoid Hub Conference on the missing shared context layer for robots — broodsugar · 2026-10-11
- NUS MAGIC Lab hiring: postdocs, PhD students and robotics researchers in Singapore — DJiafei · 2026-10-11
- StructureGS-SLAM: structure-aware Gaussian Splatting SLAM with planar instances — rsasaki0109 · 2026-10-11
- US Army field-evaluates humanoids for high-risk missions; IHMC's Alex is the only full platform winner — CyberRobooo · 2026-10-11
- Dev ports an appliance-control app to smart glasses in under 15 minutes with Codex and Lens Studio — Scobleizer · 2026-10-11
- AI gadget dot ships with Blender and Godot preinstalled — kieranklaassen · 2026-10-11