π0.5 tops the robot-policy benchmark, but task-level results still matter

chris_j_paxton · x · 2026-07-29

π0.5 leads a robot-policy benchmark, but the author warns the leaderboard is incomplete

A follow-up post from the same benchmark reports that physicalint’s π0/0.5 wins the overall comparison, with NVIDIA Robotics’ GR00T N1.7 close behind.

Other methods — Diffusion, ACT, MolmoAct 2, and SmolVLA — are described as competitive in parts of the benchmark, especially among the non-VLA policies trained from scratch. The key point is that the headline ranking alone does not capture the full picture of policy quality across tasks.

Related event: Robot Policy Benchmark Reveals Flaws in Aggregate Scoring(2 posts)→

Original post →

More from Embodied

Embodied channel →