GPT-6 Astra benchmarks revealed: 99.9% on ARC-AGI-3, 1.9x faster than GPT-5.6 Sol in Codex
whoiskatrin · x · 2026-09-04
OpenAI detailed its new flagship GPT-6 Astra with standout benchmarks and features:
- Scores 99.9% on ARC-AGI-3 and 98% on FrontierMath Tier 4; already helped solve long-standing open math problems
- State-of-the-art computer use and software engineering; produces polished documents, spreadsheets and presentations matching user templates
- 1.9x faster than GPT-5.6 Sol on Mind2Web thanks to improved Codex harness
- In Codex, it can ask questions while working independently; an experimental context feature keeps notes and searches earlier context windows across long tasks
Related event: GPT-6 Astra Scores 99.9% on ARC-AGI-3 and Builds 3D Cities in Unity(3 posts)→
More from coding & agent
- Google clarifies Antigravity terms change only covers Antigravity and Gemini CLI accounts — rseroter · 2026-09-04
- Largest open experiment finds dev tools are always mentioned but never chosen by coding agents — ycombinator · 2026-09-04
- Enterprise AI Agent Cohort Kicks Off With Engineers From PlayStation, Salesforce, ServiceNow — hugobowne · 2026-09-04
- Mathematician tests GPT-6 Astra: live Lean proof verification while writing arguments — teortaxesTex · 2026-09-04
- Makepad flow: a Rust single-executable ComfyUI alternative built in one day — anselm · 2026-09-04
- antirez: judge new models by whether they fix real blocking bugs, not three.js demos — antirez · 2026-09-04