GPT-6 Astra scores 95% on robot control task at 43% of prior cost
A set of figures circulating widely on X (Twitter) claims GPT-6 Astra scored 95% on a robot control task versus just 40% for the previous-generation Fable 5.1; it also cut output tokens by 6.2x and reduced cost by 2.3x (about 43% of the prior generation).
Confirmed
- The thread cited by @chooijeq first gave these numbers: 95% on a robot control task (Fable 5.1: 40%), 6.2x fewer output tokens, and 2.3x lower cost.
- @willcb, @AaronBergman18, @maxpaperclips, @chrisjpaxton and many others reshared the same figures, forming a propagation chain with fully consistent numbers.
Unconfirmed
- The specific benchmark, evaluation setup, and source of the 'robot control task' are not documented in primary materials; the 95% and cost figures come only from secondhand threads, with no official or paper corroboration so far.
Why it matters
- If the numbers hold up, a general chat model would substantially beat the previous dedicated version at robot control, with inference cost cut nearly in half—significant for model selection in embodied AI.
- @AaronBergman18 remarked that 'robot control got stuffed into a chat Transformer too'; in @chrisjpaxton's reshares, DJiafei quipped that instead of GPT-6 you could just call the MolmoAct2 checkpoint via API, while in @maxpaperclips' reshares Dorialexander used it to respond to an earlier European debate about LLMs—showing clear community fault lines over 'generalist vs. specialized models.'
2026-09-05 ~ 2026-09-06 · 5 related posts
- Episode 1: OpenAI launches GPT-6 'Astra' amid AGI-era claims and wave of hands-on tests(2026-09-04, 19 posts)
- Episode 2: GPT-6 Astra finds up to 176x code speedups in five minutes(2026-09-05, 2 posts)
- Episode 3: OpenAI's Astra Reportedly Trained on Over 100,000 GPUs(2026-09-05, 2 posts)
- Episode 4: GPT-6 Astra scores 95% on robot control task at 43% of prior cost(2026-09-05, 5 posts)
- Episode 5: GPT-6-Astra tops MathArena leaderboard(2026-09-06, 2 posts)
- Episode 6: Leaked Benchmarks Claim GPT-6 Astra Aces Enterprise Tasks(2026-09-06, 2 posts)
- Episode 7: Developer Says OpenAI's Astra Is First Model to Make Real Progress on His Ultra-Complex Project(2026-09-06, 2 posts)
- Episode 8: GPT-6 Astra Reportedly Beats Portal Fully Autonomously(2026-09-06, 2 posts)
Primary sources
- [source] GPT-6 Astra Scores 95% on Robot Control, 6.2x Fewer Tokens Than Fable 5.1's 40% — scaling01 · 2026-09-05
- GPT-6 hits 95% on robot control; joker asks if it just API-calls MolmoAct2 — chris_j_paxton · 2026-09-06
3 near-duplicate retellings: AaronBergman18 · willcb · max_paperclips