GPT-6 Astra scores 35.2% on ARC-AGI-3 with reasoning set to 'none', 4.5x GPT-5.6's max
mhmazur · x · 2026-09-09
ARC's mhmazur shared GPT-6 Astra's ARC-AGI-3 results, highlighting behavior with reasoning effort set to "none":
- Standard harness (model decides which notes to carry forward): Astra scored 35.2% at "none" — roughly 2x its 17.5% at "low" and close to 38.6% at "medium".
- For comparison, GPT-5.6 Sol (released July 9) scored 7.8% at "max"; Astra, eight weeks later, scored over 4.5x higher with no reasoning.
- New Provider Adapter harness (preserves opaque reasoning): 96.7% at "none", near the 98.0%–99.9% at other levels.
"None" was available in early testing but is no longer exposed via the API.
Related event: OpenAI's Astra Tops ARC-AGI-3 Benchmark(2 posts)→
More from Models
- OpenAI Images V2.5 jailbroken, guardrails bypassed — flowersslop · 2026-09-09
- NVIDIA Ships Qwen 3.8 27B NVFP4 Quantized Model on Hugging Face — TheZachMueller · 2026-09-09
- Hands-on: ChatGPT Images 2.5 shows precise local editing of infographics — xiaohu · 2026-09-09
- Early user: GPT-6 Astra's video understanding identifies everything 'with 100% accuracy' — imjustnewatai · 2026-09-09
- Should AI agents pause to think in game benchmarks? Latency vs. decision-making — imjustnewatai · 2026-09-09
- OpenAI Launches GPT-Image 2.5 With Better Multi-Turn Editing and Reference Preservation — Simon Willison · 2026-09-09