Robot Policy Benchmark Reveals Flaws in Aggregate Scoring

A robot policy benchmark spanning 12 tasks and thousands of evaluations shows π0.5 leading overall, but researchers warn that aggregate scores obscure critical task-level performance differences.

2026-07-29 ~ 2026-07-29 · 2 related posts