LIBERO-MAX Benchmark: Mid-Task World Changes Cut Every Robot Policy's Success by 11-26 Points

机器之心 · wechat · 2026-10-11

Researchers from Stanford, CMU, MIT, UIUC, and UT Austin introduced LIBERO-MAX, a benchmark testing whether robot policies adapt when the world changes mid-task. Accepted as an oral at the NeurIPS'26 Robotics World Modeling Workshop.

It uses paired rollouts — identical until a key moment, then one continues normally (Base) while the other faces one of 8 mid-execution perturbations (Dynamic) — across 8,000 paired cases, distinguishing pre-existing failures from change-induced ones. Evaluating 14 policies (VLAs, hybrids, world-action models) over 224,000 simulations, every policy dropped 11.0–25.7 points; the best (MolmoAct2, π₀.₅) still fell to 66%. Object/container moves and camera changes hurt most, and more frequent re-observation didn't close the gap.

A targeted fix — adding an image-quality detection and restoration module before X-VLA — lifted post-change success from 40.7% to 47.7%. Code, fixed case lists, and an 800-case Lite version are open-sourced.

Original post →

More from Embodied

Embodied channel →