Stanford's Mulligan targets failure states for on-robot learning, beats uniform collection across 2,550 blind evals

ZhaoMandi · x · 2026-10-10

Stanford researchers (Shuran Song, Chelsea Finn et al.) present Mulligan, a performance-guided data collection method for on-robot learning in supervised deployment. Each round, the robot runs real high-precision manipulation tasks and a human operator intervenes on failures; the next round concentrates data collection on initial states where the policy still fails, since failures concentrate in a small subset as training progresses. Across 3 real tasks and 2,550 strictly blind evaluations, it improves on uniform collection at the same operator budget. Paper, code, and data are public.

Original post →

More from Embodied

Embodied channel →