π0.5 finetuning on real manufacturing: 1 hour of clean data beat the previous 17
DominiqueCAPaul · x · 2026-09-24
A two-month ablation of π0.5 finetuning on a real manufacturing task, with all data and runs published. Final policy success rate: 98%.
Key findings:
- Pure scaling was the weakest lever: 5x more data only moved success from 63% to 76%.
- Diversity paid: 4h spread across five scenes beat 4h in the eval scene by 30pp.
- Quality crushed quantity: adding 1 clean, slower hour on top of 21h went 76% → 90% — that single hour added more than the previous 17.
- Quality alone: 1h clean data + 240 human-intervention rollouts (1.7h total) went 28% → 88%, beating the 21h model by 12 points with a twelfth of the data.
- Inference settings added more without retraining: 21h model 76% → 93%; 21h+1h model 90% → 98%.
Surprisingly unimportant: learning rate, batch size, relative vs. absolute joints, image augmentation.
More from Embodied
- Humanoid robot soccer demo: 140 simulated years of self-play training — deepakpathak · 2026-09-24
- The world is the installed base: the underrated case for humanoid robots — r0ck3t23 · 2026-09-24
- GTSAM 4.3 ships with legged navigation, CUDA-accelerated factor graphs and continuous-time trajectory estimation — fdellaert · 2026-09-24
- Workbench open-sources an agentic lab workspace for hardware engineering on the Ambion kernel — andreisavu · 2026-09-24
- Hugging Face Tutorial: Accelerating Robotics Simulation with NVIDIA Warp and MjWarp — Hugging Face Blog · 2026-09-24
- Skild AI trains Unitree G1 to play soccer with 140 years of simulated self-play — Distinct-Question-16 · 2026-09-24