Robot Leaderboard Under Fire: 5 Trials Per Task, Authors Commit to 15

YuXiang_IRVL · x · 2026-09-21

A real-robot leaderboard sparked a methodology debate. Critic Will argued 5 trials per task is a "highlight reel" lacking confidence intervals, cycle time, jam recovery, and intervention data. UCSD's YuXiang clarified the setup (5 trials × 10 tasks), committed to 15 trials per task, and noted 6-8 hours per policy across 7 policies already evaluated — underscoring the cost of rigorous real-world evaluation.

Original post →

More from Embodied

Embodied channel →