Ex-OpenAI Scientist Trains Robots with RLHF Approach: Single Policy Masters Walking, Jumping, and Crawling
chris_j_paxton · x · 2026-07-25
The author argues that wheeled robots will fundamentally lose in many environments over a 5-10 year timeline, as legged robots offer far more ways to adjust their posture and remain stable.
The quoted post demonstrates a robotics breakthrough: when physically restrained, a robot trained with a single general policy can autonomously find its way back to stable walking without any scripted recovery moves. The startup behind this was founded by a former OpenAI scientist who worked on the RLHF method powering ChatGPT, applying similar reinforcement learning concepts to robotics.
Related event: Chinese Robot Team Demonstrates Single-Policy Locomotion(2 posts)→
More from Embodied
- MouthPad^ shows how hands-free assistive hardware is being used in daily life — plopesresearch · 2026-07-25
- Autonomous fighting robot tries to flee a cage fight, and the reply is pure meme fuel — chris_j_paxton · 2026-07-25
- Eyecandy is building entertainment robots with 10–15 people and $500k in funding — retr0jirachi · 2026-07-25
- Eyecandy Robotics plans a $200 palm-sized tabletop robot for 2027 — k7agar · 2026-07-25
- An $8 ESP32-S3 now runs a fully offline, real-time object recognition camera — Yamapama · 2026-07-25
- SpaceX says Starship V3 is improving reusable orbital heat-shield tiles flight by flight — XFreeze · 2026-07-25