π0.5 tops the robot-policy benchmark, but task-level results still matter
chris_j_paxton · x · 2026-07-29
π0.5 leads a robot-policy benchmark, but the author warns the leaderboard is incomplete
A follow-up post from the same benchmark reports that physicalint’s π0/0.5 wins the overall comparison, with NVIDIA Robotics’ GR00T N1.7 close behind.
Other methods — Diffusion, ACT, MolmoAct 2, and SmolVLA — are described as competitive in parts of the benchmark, especially among the non-VLA policies trained from scratch. The key point is that the headline ranking alone does not capture the full picture of policy quality across tasks.
Related event: Robot Policy Benchmark Reveals Flaws in Aggregate Scoring(2 posts)→
More from Embodied
- Unitree G1 appears to hop out of a car in a striking robot demo — chris_j_paxton · 2026-07-29
- Tesla claims one Robotaxi equals seven Model Ys in economic output — JOBhakdi · 2026-07-29
- World Labs argues the key robotics metric is sim-to-real success rate alignment — andrew_n_carr · 2026-07-29
- Interactive rings may need haptic output, not just voice input, for AI use — plopesresearch · 2026-07-29
- Hackathon team uses OpenAI Codex to drive a robot arm from natural language — davidfromkansas · 2026-07-29
- Bagel Labs releases WorldDiT, a sub-billion-parameter robot control model — EXM7777 · 2026-07-29