Two months of π0.5 finetune ablations on a real manufacturing task: data quality beats volume
DominiqueCAPaul · x · 2026-09-24
DominiqueCAPaul published all results, data, and runs from two months of ablating π0.5 finetunes on a real manufacturing task, reaching a 98% success policy.
Key findings:
- Pure scaling was the weakest lever: 5x more data moved success from 63% to 76%.
- Diversity mattered more: 4h spread across five scenes beat 4h in the eval scene by 30pp.
- Quality was strongest: one clean, slower hour on top of 21h pushed 76% → 90%—more than the previous 17 hours combined.
- Quality alone worked too: 1h of clean data plus 240 human-intervention rollouts (1.7h total) went 28% → 88%.
Takeaway: in robot learning, data quality and diversity can matter far more than sheer volume.
Related event: π0.5 Fine-Tuning Study: 1 Hour of Clean Data Beats 17 Hours of Scale(6 posts)→
More from Embodied
- Schmidhuber: No AI robot can match a plumber — screen-world AI isn't true intelligence — SchmidhuberAI · 2026-09-24
- The world is the installed base: the underrated case for humanoid robots — r0ck3t23 · 2026-09-24
- GTSAM 4.3 ships with legged navigation, CUDA-accelerated factor graphs and continuous-time trajectory estimation — fdellaert · 2026-09-24
- Workbench open-sources an agentic lab workspace for hardware engineering on the Ambion kernel — andreisavu · 2026-09-24
- Hugging Face Tutorial: Accelerating Robotics Simulation with NVIDIA Warp and MjWarp — Hugging Face Blog · 2026-09-24
- Skild AI trains Unitree G1 to play soccer with 140 years of simulated self-play — Distinct-Question-16 · 2026-09-24