Policies trained fully inside world models deploy zero-shot on real robots, beating PPO

chris_j_paxton · x · 2026-09-04

A new preprint by Joseph Amigo, Rooholla Khorrambakht, Nicolas Mansard, and Ludovic Righetti trains RL policies entirely inside learned world models — no simulator — and deploys them zero-shot on real robots.

The problem: model-based RL in world models is usually intractable, since accurate gradients require capturing contact dynamics with a very large model.

Method: a decoupled first-order gradient (FoG) RL scheme —

Results: zero-shot real-world deployment on a tabletop manipulator (Push-T), a Unitree G1 humanoid (ego-centric grasp and lift), and a Go2 quadruped (push cube). On the canonical real-world Push-T benchmark, sample efficiency is significantly better than PPO, with similar gains on ego-centric manipulation.

Related event: Robots learn contact-rich manipulation in world models without simulators(2 posts)→

Original post →

More from Embodied

Embodied channel →