Play2Perfect (CoRL 2026): Play Pretraining Yields Precise Zero-Shot Robot Assembly

leto__jean · x · 2026-09-11

Stanford's Play2Perfect, accepted to CoRL 2026, uses a 2-stage RL pipeline: free-space "play" pretraining to build manipulation priors, then sparse-reward finetuning on contact-rich precise assembly (screwing, tight insertion, real-size plug/fork), with zero-shot sim-to-sim transfer from Isaac Sim to MuJoCo. Systematic ablations show pretraining transfers best when robots manipulate objects in-hand with fingers; from-scratch training with hand-crafted dense rewards stalls near zero. The team also generated 2 new envs from scratch with Opus 5, reusing the same RL code. An interactive browser demo is live.

Related event: Stanford's Play2Perfect accepted to CoRL 2026 with zero-shot assembly(2 posts)→

Original post →

More from Embodied

Embodied channel →