Robot Leaderboard Under Fire: 5 Trials Per Task, Authors Commit to 15
YuXiang_IRVL · x · 2026-09-21
A real-robot leaderboard sparked a methodology debate. Critic Will argued 5 trials per task is a "highlight reel" lacking confidence intervals, cycle time, jam recovery, and intervention data. UCSD's YuXiang clarified the setup (5 trials × 10 tasks), committed to 15 trials per task, and noted 6-8 hours per policy across 7 policies already evaluated — underscoring the cost of rigorous real-world evaluation.
More from Embodied
- Apple's MintAct unifies GUI agents across mobile, desktop, and web, hitting SOTA 48.9 on OSWorld-Verified — apple · 2026-09-21
- Movement Trend Guidance lifts 3D diffusion policies without explicit trajectories — 72% vs 49% on real robots — dalian-university-of-technology · 2026-09-21
- OpenAI robotics hiring surges as company re-enters the humanoid race five years after shutdown — ocean_protocol · 2026-09-21
- OpenAI robotics job listings jump from 11 to 27, base salaries up to $500,000 — Dr_Singularity · 2026-09-21
- Tendon-Driven Robot Hand Demo Shows High-DOF, High-Speed Dexterous Manipulation — LexiLove · 2026-09-21
- 1X targets 50,000 humanoid robots shipped by 2027, ramping to 110k/year capacity — Scobleizer · 2026-09-21