GPT-6 Astra robot evals show strong task decisions but weak physical control
Galbot · hf · 2026-10-01
Galbot systematically evaluated GPT-6 Astra as a general-purpose embodied policy across six domains:
- Gripper manipulation: corrects task targets and preps contact conditions; hybrid control with π0.5 hits 48% success on a RoboDojo subset.
- Dexterous manipulation: 50% success across ten DexJoCo trials with hybrid control, but direct in-hand control struggles to coordinate finger contacts.
- Mobile manipulation: 38.7% success on RoboCasa365.
- Navigation: leads local comparisons with 92% on RxR instruction following and 82% on HM3D object search, though search involves substantial detours.
- Locomotion: unreliable — none of five sequential obstacle-course attempts reached the goal.
- Humanoid loco-manipulation: beats baselines on 13 of 30 HumanoidBench tasks.
The core finding: a gap between useful task decisions and reliable physical control. Inference latency is a major constraint — policy-assisted and direct control consumed 624.8M and 1.132B tokens respectively across conditions, and a 30-second locomotion run required 250 model calls averaging 39.86 seconds each, with physics paused during inference.
More from Embodied
- Humanoid robot demo shows real-time perception, hand coordination and human interaction — CurieuxExplorer · 2026-10-01
- Robotics companies aren't taking 'cheapness' seriously enough, argues founder — hudzah · 2026-10-01
- Blogger predicts humanoid robots will soon run hotel services like laundry — CyberRobooo · 2026-10-01
- NVIDIA's Instant NuRec reconstructs a drivable 3DGS world from driving logs in ~1.5 seconds — rsasaki0109 · 2026-10-01
- Yuanyjie Robotics raises seed round to automate restaurant kitchens with embodied AI — 创业邦 · 2026-10-01
- RoboCoach uses world-model-imagined failures as coaching: 13.3% to 75% success with 150 demos — Tsinghua · 2026-10-01