Robot-policy benchmark says overall scores hide more than they reveal

chris_j_paxton · x · 2026-07-29

A robot-policy benchmark says evaluation is broken

A thread about a study that ran thousands of evaluations across 12 manipulation tasks to compare robot policies.

The benchmark’s takeaway is blunt: simple aggregate scores can be misleading. In the related results table, π0.5 comes out ahead overall, with GR00T N1.7 close behind, while Diffusion, ACT, MolmoAct, and SmolVLA trail on the headline metric. But the author says the overall ranking does not tell the full story, implying task-level behavior and robustness matter too.

Related event: Robot Policy Benchmark Reveals Flaws in Aggregate Scoring(2 posts)→

Original post →

More from Embodied

Embodied channel →