GPT-6 Astra as robot policy: 22.48% success, beats every public RoboDojo entry
ben_j_todd · x · 2026-09-22
A new arXiv paper evaluates "LLM as policy": using LLMs directly for robot manipulation without task-specific finetuning. Across all 42 RoboDojo tasks and 2,100 trials, GPT-6 Astra achieved a 22.48% average success rate (28.97 Score), ranking above all 40 public policies, while GPT-5.5 and DeepSeek-Flash managed only 0.88% and 1.92% with the same post-processing. Astra's profile is sharply polarized: strong on semantically-demanding tasks, weak on precision, dynamic control, and complex bimanual coordination. One-shot demos showed no aggregate benefit, though traces reveal within-episode self-corrections under perturbation.
More from Embodied
- RoboDawn: Tsinghua & Tencent Hunyuan drive robots with frozen VLM, 73.6% one-shot success — _akhaliq · 2026-09-22
- PragmaBot: robots learn online from real-world failures without retraining, RA-L/IROS 2026 — ChongZzZhang · 2026-09-22
- Humanoid demand could hit hundreds of millions; millions built by 2029, says Robostrategy — Rewkang · 2026-09-22
- Intrinsic open-sources Intrinsic Core: ROS-compatible building blocks for physical AI robotics — facontidavide · 2026-09-22
- MolmoAct2 gets out-of-the-box support for SO-101 arms and YAM, plus VR teleop — k7agar · 2026-09-22
- Quest-based robot data collection adds passthrough and headless modes, YAM teleop switching — k7agar · 2026-09-22