OpenAI's Astra scores 99.9% on ARC-AGI-3, far surpassing prior models' ~30%

downingARK · x · 2026-09-09

OpenAI's newly released Astra model all but maxed out the ARC-AGI-3 reasoning benchmark with a score of 99.9%, compared with scores in the 30s for prior GPT-5.6 Sol and Opus 5.

Astra's computer-use capabilities are also reportedly 2× faster than previous models, which could drive more usage of the ChatGPT app.

Related event: OpenAI's Astra Tops ARC-AGI-3 Benchmark(2 posts)→

Original post →

More from Models

Models channel →