Policies trained fully inside world models deploy zero-shot on real robots, beating PPO
chris_j_paxton · x · 2026-09-04
A new preprint by Joseph Amigo, Rooholla Khorrambakht, Nicolas Mansard, and Ludovic Righetti trains RL policies entirely inside learned world models — no simulator — and deploys them zero-shot on real robots.
The problem: model-based RL in world models is usually intractable, since accurate gradients require capturing contact dynamics with a very large model.
Method: a decoupled first-order gradient (FoG) RL scheme —
- a large global world model (DIAMOND for Push-T, DreamerV4 for ego-centric tasks) generates high-fidelity forward trajectories;
- a lightweight latent-space surrogate approximates its local dynamics for efficient gradient computation;
- both the world model and policies are learned from scratch from real-world data.
Results: zero-shot real-world deployment on a tabletop manipulator (Push-T), a Unitree G1 humanoid (ego-centric grasp and lift), and a Go2 quadruped (push cube). On the canonical real-world Push-T benchmark, sample efficiency is significantly better than PPO, with similar gains on ego-centric manipulation.
Related event: Robots learn contact-rich manipulation in world models without simulators(2 posts)→
More from Embodied
- Smart ring packs a microphone, heart rate, SpO2 sensors and vibration motor into one — AnhPhuNguyen1 · 2026-09-04
- Astra: More Aligned but Less Monitorable? Podcast Digs Into the Tradeoff — The Cognitive Revolution · 2026-09-04
- Indie dev nears pre-orders for an AI handheld for kids blending stories and learning — pramodk73 · 2026-09-04
- Free 39-Episode Control Bootcamp: The Control Theory That Runs Real Robots — lukas_m_ziegler · 2026-09-04
- AR glasses as robot navigation interface: open-source Unitree Go2 project wins Lenslist — Scobleizer · 2026-09-04
- DLSS 5 ships with NBA 2K27: costs half your 4K frame rate, worth it — ryanshrout · 2026-09-04