GPT-6 Astra as robot policy: 22.48% success, beats every public RoboDojo entry

ben_j_todd · x · 2026-09-22

A new arXiv paper evaluates "LLM as policy": using LLMs directly for robot manipulation without task-specific finetuning. Across all 42 RoboDojo tasks and 2,100 trials, GPT-6 Astra achieved a 22.48% average success rate (28.97 Score), ranking above all 40 public policies, while GPT-5.5 and DeepSeek-Flash managed only 0.88% and 1.92% with the same post-processing. Astra's profile is sharply polarized: strong on semantically-demanding tasks, weak on precision, dynamic control, and complex bimanual coordination. One-shot demos showed no aggregate benefit, though traces reveal within-episode self-corrections under perturbation.

Original post →

More from Embodied

Embodied channel →