Stanford's Mulligan targets failure states for on-robot learning, beats uniform collection across 2,550 blind evals
ZhaoMandi · x · 2026-10-10
Stanford researchers (Shuran Song, Chelsea Finn et al.) present Mulligan, a performance-guided data collection method for on-robot learning in supervised deployment. Each round, the robot runs real high-precision manipulation tasks and a human operator intervenes on failures; the next round concentrates data collection on initial states where the policy still fails, since failures concentrate in a small subset as training progresses. Across 3 real tasks and 2,550 strictly blind evaluations, it improves on uniform collection at the same operator budget. Paper, code, and data are public.
More from Embodied
- Starkey Unveils Omega AI+ Hearing Aids With G4 Gen AI Processor and Quad DNN — BrandonSawalich · 2026-10-10
- Wayve CEO demos AI Driver with Stellantis in a Fiat 500e on Turin's tight streets — alexgkendall · 2026-10-10
- 192GB Framework Desktop Batch 2 sells out; 128GB still in stock — film_girl · 2026-10-10
- Danu Robotics' six-year fight to build a better recycling robot — TechCrunch AI · 2026-10-10
- NVIDIA releases GR00T-based Agile One S SSD Pick robot deployment model on Hugging Face — _akhaliq · 2026-10-10
- Success-Guided Sampling: sim-to-real RL nails dexterous assembly with zero demos — kevin_zakka · 2026-10-10