GPT-6 Astra robot evals show strong task decisions but weak physical control

Galbot · hf · 2026-10-01

Galbot systematically evaluated GPT-6 Astra as a general-purpose embodied policy across six domains:

The core finding: a gap between useful task decisions and reliable physical control. Inference latency is a major constraint — policy-assisted and direct control consumed 624.8M and 1.132B tokens respectively across conditions, and a 30-second locomotion run required 250 model calls averaging 39.86 seconds each, with physics paused during inference.

Original post →

More from Embodied

Embodied channel →