Robotics' real frontier is post-training: self-play in sim closes the demo-to-deployment gap

ZGojcic · x · 2026-09-24

Abhinav Gupta's team introduces a self-play approach to post-train robotic policies in simulation, emerging robust behaviors before hardware deployment. Their argument: while attention goes to pre-training, post-training is what closes the gap between a demo and deployment—"between a video and actual dollars"—making it the real frontier in robotics. Inspired by DeepMind's original self-play work and their own robust adversarial RL research. Commenters add that labs that crack sim post-training will lap everyone.

Original post →

More from Embodied

Embodied channel →