Robotics benchmarks may reward 0.2% gains while real-world success collapses from 92% to 43%
kscottz · x · 2026-07-28
A quote about robotics argues that benchmarks are rewarding teams for squeezing out 0.2% better scores, which encourages overfitting instead of better datasets.
The post says the real problem is not model micromanagement but the boring work of building data that matches deployment environments. It gives a concrete example: robots keep being evaluated folding T-shirts on tables, not doing actual laundry work in laundromats.
The quoted thread also points to a harsh real-world gap in robotics evaluation: one team saw 92% success in internal tests but only 43% at the customer site.
More from Embodied
- A robotaxi spotted in Los Angeles sparks questions about paid rides — mrjonfinger · 2026-07-28
- Japan industrial robot orders rose 40% to a record ¥312.7 billion as AI spending accelerated — 创业邦 · 2026-07-28
- Survey maps progress reward modeling across robotic learning and benchmarks — northwestern-university · 2026-07-28
- Embodied manipulation gets a five-layer data pyramid for robot alignment — PekingUniversity · 2026-07-28
- Dex-Net in 60 Seconds: A Classic Look at Robotic Grasping Uncertainty — berkeley_ai · 2026-07-28
- Being-H0.8 adds tactile prediction to robot world models with 500,000 hours of video — 量子位 · 2026-07-28