GPT-6 Astra hits 62.7% on ARC-AGI-3, 99.9% with new adapter harness, sets ARC-AGI-2 SOTA at 95%
BlackHC · x · 2026-09-04
ARC Prize's analysis of OpenAI's GPT-6 Astra shows record-breaking results across the ARC-AGI suite:
- ARC-AGI-3: 62.7% under the standard harness (models carry notes forward, apples-to-apples with history), more than doubling the previous verified high score and surpassing human performance on 96% of levels
- New provider adapter harness: preserving the model's opaque reasoning state between requests with native compaction, Astra scores 99.9% on ARC-AGI-3; ARC Prize calls it the most precise symbolic model of novel environments they've seen
- ARC-AGI-2: new SOTA of 95.0%; ties Fable 5 at 98.5% on ARC-AGI-1
- First backward bend in the cost/performance curve, with max run cost at $26K
ARC Prize will now evaluate models with both harnesses wherever possible.
More from Models
- Tavus Launches Sparrow-2, Claiming #1 in End-of-Turn Detection and Interruption Handling — ycombinator · 2026-09-04
- ARC-AGI-3: 10x reasoning tokens cuts total cost from $48k to $26k vs medium — i_dg23 · 2026-09-04
- Researchers flag data contamination concerns in benchmark behind Astra's time-horizon score — dfrsrchtwts · 2026-09-04
- antirez: judge new models by whether they fix real blocking bugs, not three.js demos — antirez · 2026-09-04
- Astra tops ARC-AGI-3 as the most cost-efficient entry on the leaderboard — flowersslop · 2026-09-04
- GPT-6 Astra shows near 10x jump in no-CoT capability, making it far less monitorable — rickasaurus · 2026-09-04