GPT-6 Astra sparks AGI debate: 99.9% vs 62.7% ARC-AGI-3 scores explained
johnseach · x · 2026-09-06
OpenAI shipped GPT-6 Astra on September 3, and Greg Brockman told reporters "Welcome to the AGI era" — but the numbers deserve fine print.
- Capabilities: Astra is a strong tool-using agent, far better at driving computers than GPT-5.6 Sol; it works through forms, browsers, CRMs, calendars, and real software like KiCad, Unity, and CAD; 98% on FrontierMath Tier 4 and the first OpenAI model to hit the internal "Critical" cyber threshold (public version still refuses advanced exploit work)
- Specs/pricing: 1.05M token context, API near $10/$50 per million input/output tokens, trusted orgs and paid plans first
- The benchmark catch: the viral 99.9% ARC-AGI-3 score used OpenAI's own Provider Adapter harness; ARC Prize's standard harness put the same model at 62.7% at max effort — both real, measuring different systems
- AGI verdict: no agreed test; Altman himself called AGI a poorly defined marketing term; Claude Fable 5.1 and other frontier models sit close on coding and agent indexes
Treat it as a real jump in agentic computer work, not a scientific AGI result.
Related event: OpenAI Launches GPT-6 Astra, Smashing Benchmarks and Igniting AGI Debate(8 posts)→
More from Models
- Users allege Astra's Codex quota accounting consumes 4-5x Sol's allowance, not 2.5x — Medical-Yam3367 · 2026-09-06
- Astra scores 77.3% on Browser Use Benchmark v2, crushing Opus 5's 50.5% — gabrielchua · 2026-09-06
- Astra Ultra with board visualization loses to 1800-ELO chess bot after beating 1500 — MikePFrank · 2026-09-06
- First impression: Astra claims PTX restriction bypassable, but Sol was right — A_K_Nain · 2026-09-06
- Insider teases that next week's demos will far outshine OpenAI's official blog and trailer — ChrisGPT · 2026-09-06
- GPT-6 Astra system card: first OpenAI model to hit Critical cybersecurity threshold — RyanGreenblatt · 2026-09-06