Two months of π0.5 finetune ablations: 1 hour of clean data beats 17 hours of scale
DominiqueCAPaul · x · 2026-09-24
The author spent two months ablating π0.5 finetunes on a real manufacturing task, reaching a 98% policy success rate, and will publish all results, data, and runs. Key findings: 5x more data was the weakest lever (63%→76%); spreading 4h across five scenes beat 4h in the eval scene by 30pp; adding just 1 clean hour on top of 21h jumped 76%→90% — more than the previous 17 hours; and 1h clean data plus 240 human-intervention rollouts (1.7h total) lifted 28%→88%. Data quality, not scale, is the lever.
Related event: π0.5 Fine-Tuning Study: 1 Hour of Clean Data Beats 17 Hours of Scale(6 posts)→
More from Embodied
- Google's DynamicWebPaige: the next decade is software that picks things up — DynamicWebPaige · 2026-09-24
- Microsoft's first new Surface Mouse in 10 years adds haptics and a $79.99 Copilot button — tomwarren · 2026-09-24
- Microsoft refreshes small Surface Pro and Laptop with Snapdragon X2 Plus, prices jump to $1,149+ — tomwarren · 2026-09-24
- Training real hardware, not a digital twin: learning directly in nonlinear wave systems — bravo_abad · 2026-09-24
- Schmidhuber: No AI robot can match a plumber — screen-world AI isn't true intelligence — SchmidhuberAI · 2026-09-24
- The world is the installed base: the underrated case for humanoid robots — r0ck3t23 · 2026-09-24