LIBERO-MAX: mid-task changes cut success rates for all 14 robot policies by 11-25.7 points

DJiafei · x · 2026-10-08

Researchers released LIBERO-MAX, a dynamic robustness benchmark testing whether robot policies can still finish tasks when the world changes mid-execution — e.g., the cup moves after the arm commits.

Across all 14 tested robot policies, success drops by 11.0–25.7 percentage points. Code, dataset, and paper are public.

Original post →

More from Embodied

Embodied channel →