OpenAI's Astra scores 99.9% on ARC-AGI-3, far surpassing prior models' ~30%
downingARK · x · 2026-09-09
OpenAI's newly released Astra model all but maxed out the ARC-AGI-3 reasoning benchmark with a score of 99.9%, compared with scores in the 30s for prior GPT-5.6 Sol and Opus 5.
Astra's computer-use capabilities are also reportedly 2× faster than previous models, which could drive more usage of the ChatGPT app.
Related event: OpenAI's Astra Tops ARC-AGI-3 Benchmark(2 posts)→
More from Models
- Epoch AI: GPT's Quadratic Latency vs Claude's Linear May Explain Pricing Gap — scaling01 · 2026-09-09
- Hackers are stealing Claude tokens from subscribers, Anthropic warns — TechCrunch AI · 2026-09-09
- Grok on AI cracking Navier-Stokes: not surprising, AI stood on giants' shoulders — robleclerc · 2026-09-09
- GPT-6 Astra tops Andon Labs' Vending-Bench, unseating Fable — Outside-Iron-8242 · 2026-09-09
- OpenAI: training new internal model with 'unprecedented' math benchmark performance — TheZvi · 2026-09-09
- DeepSeek v4.1 Flash flash sale: 58M tokens for $1, available for 2 days only — MicahBerkley · 2026-09-09