A robotics policy trained on 4 hours of diverse data beats in-domain data

DominiqueCAPaul · x · 2026-07-27

The author reports an unexpected robotics finding: a policy trained on 4 hours of data gathered across 5 tables and 2 workstations outperformed, on the lightbox evaluation, a policy trained on the same amount of data collected inside the lightbox itself.

The main takeaway is that diverse out-of-domain data beat in-domain data on an in-domain benchmark. The author says they expected hyperparameters to matter more, but data collection strategy turned out to be the bigger factor.

Original post →

More from Embodied

Embodied channel →