Ex-OpenAI Scientist Trains Robots with RLHF Approach: Single Policy Masters Walking, Jumping, and Crawling

chris_j_paxton · x · 2026-07-25

The author argues that wheeled robots will fundamentally lose in many environments over a 5-10 year timeline, as legged robots offer far more ways to adjust their posture and remain stable.

The quoted post demonstrates a robotics breakthrough: when physically restrained, a robot trained with a single general policy can autonomously find its way back to stable walking without any scripted recovery moves. The startup behind this was founded by a former OpenAI scientist who worked on the RLHF method powering ChatGPT, applying similar reinforcement learning concepts to robotics.

Related event: Chinese Robot Team Demonstrates Single-Policy Locomotion(2 posts)→

Original post →

More from Embodied

Embodied channel →