Two months of π0.5 finetune ablations: 1 hour of clean data beats 17 hours of scale

DominiqueCAPaul · x · 2026-09-24

The author spent two months ablating π0.5 finetunes on a real manufacturing task, reaching a 98% policy success rate, and will publish all results, data, and runs. Key findings: 5x more data was the weakest lever (63%→76%); spreading 4h across five scenes beat 4h in the eval scene by 30pp; adding just 1 clean hour on top of 21h jumped 76%→90% — more than the previous 17 hours; and 1h clean data plus 240 human-intervention rollouts (1.7h total) lifted 28%→88%. Data quality, not scale, is the lever.

Related event: π0.5 Fine-Tuning Study: 1 Hour of Clean Data Beats 17 Hours of Scale(6 posts)→

Original post →

More from Embodied

Embodied channel →