GPT-6 Astra hits 99% on ARC-AGI-3 via provider adapter harness, up from 62.7%
rohanpaul_ai · x · 2026-09-05
ARC Prize announces OpenAI's GPT-6 Astra achieves SOTA on ARC-AGI: 63% on ARC-AGI-3 under the provider-neutral harness, 99% via a new provider adapter harness, surpassing human performance on 96% of levels and building the most precise symbolic model of novel environments seen to date.
The striking part: identical weights score 62.7% with ARC Prize's standard harness but 99.9% when allowed OpenAI's native reasoning-state preservation and long-context management. The benchmark designed to expose novel-reasoning failures is now near ceiling — and the evaluation wrapper alone swings results by 37 points.
Related event: GPT-6 Astra saturates ARC-AGI-3 with near-zero reasoning tokens(6 posts)→
More from Models
- ChatGPT's image generator is much improved via Astra — joshua_saxe · 2026-09-05
- Fable 5.1 vs GPT 6 Astra on 3D Blender asset generation shows a stark gap — curious_capsuleer · 2026-09-05
- GPT 6 'Astra' reportedly recreates Pokémon from a single prompt — IanArawjo · 2026-09-05
- User switches back to GPT from Gemini after 9/10 PDF failures and heavy usage burn — DrlNoV · 2026-09-05
- GPT-6 Astra's experimental compaction in Codex saves notes across context windows, off by default — TheMoonMidas · 2026-09-05
- Same prompt run on Opus 4.8, Fable 5.1, and GPT 6 Astra for comparison — LightningMcLovin · 2026-09-05