GPT-6 Astra saturates ARC-AGI-3 with near-zero reasoning tokens
GPT-6 Astra achieved an officially certified SOTA result on ARC Prize's interactive reasoning benchmark ARC-AGI-3, and its reasoning style differs notably from previous models, sparking broad community discussion of "implicit reasoning."
Confirmed
- ARC Prize officially (relayed by François Chollet) announced GPT-6 Astra achieved a "step-change" capability leap on ARC-AGI-3: 66% under the standard test harness, near-perfect (99.9%) using the continuous-conversation/provider adapter harness, with performance on 96% of ARC-AGI games (m3, m5).
- The same model's scores vary enormously across evaluation frameworks: jumping from roughly 62.7%/63% under the standard framework to 99% under the adapter harness, showing harness design significantly affects results (m3).
- A user testing under the provider-adapter harness found Astra hit 97% without any chain-of-thought (CoT); poster @haider1 called the result "pretty insane" (m2).
- Charts released by Mike Knoop (Metaculus founder, associated with ARC Prize) showed: at lower reasoning settings Astra often outputs zero reasoning tokens per action, yet its accuracy is double that of GPT-5.6 Sol (m1).
- Behavioral stats from @mhmazur: in the standard environment at low reasoning effort, 98% of GPT-5.6 Sol's actions use reasoning tokens, versus only 27% for GPT-6 Astra, hinting at a major leap in implicit reasoning (m4).
Why it matters
- Astra scores high while generating almost no explicit reasoning tokens, challenging the assumption that "more reasoning tokens = better performance" and possibly marking a new paradigm of internalized reasoning ability.
- The 30+ percentage-point gap between the standard and specialized harnesses once again underscores how critical evaluation methodology is to assessing model capability.
2026-09-04 ~ 2026-09-05 · 6 related posts
- Episode 1: GPT-6 Astra saturates ARC-AGI-3 with near-zero reasoning tokens(2026-09-04, 6 posts)
- Episode 2: GPT-6 Astra's Token Efficiency Emerges as Hidden Advantage(2026-09-05, 3 posts)
- Episode 3: Developers Call GPT-6 Astra's Computer Use 'Superhuman'(2026-09-05, 2 posts)
Primary sources
- GPT-6 Astra hits 66% on ARC-AGI-3, near-100% with custom harness at ~$360 per game — AccBalanced ·
- Astra Fully Saturates the ARC-AGI-3 Benchmark Using Fewer Moves Than the Average Human — ObiWanCanownme ·
- GPT-6 Astra uses reasoning tokens in only 27% of ARC-AGI-3 actions, hinting at a leap in opaque reasoning — mhmazur ·
- [source] Astra Fully Saturates the ARC-AGI-3 Benchmark Using Fewer Moves Than the Average Human — ObiWanCanownme · 2026-09-04
- GPT-6 Astra reportedly hits 97% on ARC-AGI-3 without CoT, even under a provider-adapter harness — haider1 · 2026-09-04
- [source] GPT-6 Astra uses reasoning tokens in only 27% of ARC-AGI-3 actions, hinting at a leap in opaque reasoning — mhmazur · 2026-09-05
- ARC v3: Astra Low Emits Zero Reasoning Tokens Yet 2x More Accurate Than Sol Max — rbhar90 · 2026-09-05
- GPT-6 Astra hits 99% on ARC-AGI-3 via provider adapter harness, up from 62.7% — rohanpaul_ai · 2026-09-05
- [source] GPT-6 Astra hits 66% on ARC-AGI-3, near-100% with custom harness at ~$360 per game — AccBalanced · 2026-09-05