GPT-6 Astra early reviews: record ARC-AGI scores and big speed gains, but benchmark gaps spark debate
OpenAI's unreleased flagship GPT-6 Astra has seen a wave of early evaluations and community discussion. The overall picture: eye-catching benchmark scores with a major jump in reasoning cost-effectiveness, but clear gaps between official numbers and independent tests—and several reviewers are more inclined to see it as a computer-use/agent model rather than AGI.
Confirmed
- ARC-AGI: ARC Prize team member mhmazur wrote a long post saying Astra scored 62.7% under the standard harness, a record; and 99.9% under the new architecture (m11). The official ARC Prize blog also noted Astra hit 99.9% on ARC-AGI-3 at lower cost than Opus 5's mere 30% a month earlier (m14). Reddit user DeArgonaut compiled all of Astra's official scores and cost tiers across ARC-AGI-1/2/3 (m8).
- Hands-on comparisons: Bindu Reddy's testing found Astra excellent but still slightly behind Fable 5.1, weak on long-running agentic loops, strong on browser operation and 3D rendering, while being much cheaper and much faster (m4, m6). tristanbob also confirmed Astra is noticeably faster than Fable, unsure whether from token speed or token efficiency, jokingly calling it "Fastra" (m12).
- Official launch figures include 97.6% on FrontierMath Tier 4, 96.0% on GPQA Diamond, and 100% on ExploitBench (m16).
- Other takes: petergostev's long hands-on write-up says Astra's intelligence jumped significantly, with one case where SQL code slimmed from 33,000 lines to 1,800 (m5, m10); MoonL88537's first impression is that it's smarter and steadier in judgment than Fable, making flawless instant calls where Sol and Fable err (m13). 潜空间's roundup notes Astra's Coding Agent capability is on par with Claude Opus 5 and Fable 5, behind Fable, with per-task cost halved but chain-of-thought monitoring broken (m15).
Unconfirmed
- Artificial Analysis's neutrality was questioned by ShubhamGarg123, who claims Astra may actually lag Fable or even Anthropic's Opus, while the outlet calls itself independent yet is cited by millions (m1).
- Simon Smith flagged an anomaly: Astra set a record on Epoch ECI yet merely tied GPT-5.6 Sol on the Artificial Analysis Intelligence Index; users question the differing methodologies (m7).
- eyishazyer broke down the official-vs-independent gap, saying the leaderboard's 97.6% corresponds to just 61 points in independent testing—double the price for the same intelligence (m16).
- VraserX's tweet claiming Astra clearly beats Fable 5.1 in a scene-generation demo and is already AGI is a one-sided enthusiast take with no systematic evaluation data (m9).
Why it matters
- Astra shows a generational leap in cost-effectiveness on abstract reasoning benchmarks (higher scores at lower cost); if independently reproduced, it will shape judgments about frontier model capability boundaries.
- The systematic gap between official numbers and independent evaluations, plus a third-party benchmark org being accused of bias, highlights the current measurement chaos in model evaluation—users should treat launch figures with caution.
- The consensus direction among several reviewers (becomingengageably and others): rather than debating whether Astra is AGI, focus on its performance in real workflows like browser operation and long-horizon agent loops—a more pragmatic framework for evaluating the new generation of models.
2026-09-04 ~ 2026-09-05 · 16 related posts
Primary sources
- GPT-6 Astra's full ARC-AGI 1/2/3 results compiled with cost-tier breakdowns — DeArgonaut · 2026-09-04
- GPT-6 Astra scores 99.9% on ARC-AGI-3 for less than Opus 5's 30% run cost — 233C · 2026-09-04
- GPT-6 Astra reviews: half the per-task cost, but CoT monitoring breaks down — vista8 · 2026-09-04
- [source] GPT-6 'Astra' Smashes ARC-AGI Records: 62.7% on Standard Harness, 99.9% on ARC-AGI-3 — repligate · 2026-09-04
- [source] Astra scores 61 on independent index despite 97.6% FrontierMath hype, at 2.5x the price — eyishazyer · 2026-09-04
- GPT-6-Astra hands-on: smarter and creatively persistent, cuts 33k lines of SQL to 1.8k — i_dg23 · 2026-09-04
- GPT-6 Astra Sets Epoch ECI Record but Matches GPT-5.6 on AA Index, Sparking Benchmark Doubts — scaling01 · 2026-09-04
- GPT-6 Astra Demolishes Fable 5.1 in Scene Generation Demo, Claims 'Astra Is AGI' — VraserX · 2026-09-04
- Rumor: AI Analysis benchmark accused of favoring OpenAI as GPT-6 Astra allegedly trails Fable and Opus — Shubham_Garg123 · 2026-09-04
- GPT-6 Astra Benchmarks Split: A Computer-Use Specialist, Not AGI — becomingengageably · 2026-09-04
- Reviewer: OpenAI's GPT-6-Astra finally 'gets what you mean,' with Fable-level intelligence and real gains in game dev — pvncher · 2026-09-05
- [source] Early tests: GPT-6 Astra slightly below Fable 5.1 but much cheaper and faster — bindureddy · 2026-09-05
- GPT-6 Astra is much cheaper and faster than Fable 5.1, but slightly behind on agentic loops — bindureddy · 2026-09-05
- First impressions of GPT-6 Astra: noticeably smarter and more sensible than Fable — MoonL88537 · 2026-09-05
- Early user hands-on: OpenAI's Astra model is noticeably faster than Fable — tristanbob · 2026-09-05
1 near-duplicate retellings: becomingengageably