Two months of π0.5 finetuning ablations: diversity beats raw scaling, 98% success
philfung · x · 2026-09-24
The team shares all ablation results from two months of finetuning π0.5 on a real manufacturing task, reaching 98% policy success and publishing every run including failures. Key takeaways: 5x data only lifted success from 63% to 76% (scaling was the weakest lever), while spreading the same 4 hours of data across five scenes beat concentrated collection—diversity extracts more from identical data budgets.
More from Embodied
- Robot Fashion Is Here: Custom Outfits for Quadrupeds After ICSR 2026 Show — heatherknight · 2026-09-24
- Same model, two placements: agentic control completes 18/20 LEGO tasks vs 6/20 for code-as-policy — paigeinsf · 2026-09-24
- After Astra: where does value go when frontier labs ship robotics brains? — mihdalal · 2026-09-24
- Robotics' GPT-3 moment may just be GPT-6 itself, argues investor — mihdalal · 2026-09-24
- Qualcomm and Prism ML Run 1-bit Bonsai VLM Locally on Snapdragon AR1 Smart Glasses — Scobleizer · 2026-09-24
- Opus 5.5 Averages Under $1 Per Robotics Trial, Scoring 1.8x Opus 5 at Half the Cost — ycombinator · 2026-09-24