GPT-6 Astra as a Quadruped Locomotion Policy: 250 Inferences, 5 Seconds of Walking

YuXiang_IRVL · x · 2026-09-09

Responding to a professor's challenge that an LLM could directly implement a quadruped locomotion policy, Srinivas ran the experiment with GPT-6 Astra: the model output joint targets at 50 Hz like an RL policy, controlling a simulated Unitree Go1 with physics paused between calls. 250 inferences produced 5 seconds of walking — a playful test of whether frontier models, if orders of magnitude faster and cheaper, could serve as edge-device control policies.

Related event: GPT-6 Astra Directly Controls Quadruped Locomotion(2 posts)→

Original post →

More from Embodied

Embodied channel →