OpenAI Launches GPT-6 Astra: 99.9% on ARC-AGI-3, but Independent Evals Call It Uneven
Latent Space · rss · 2026-09-04
- The launch: OpenAI released GPT-6 Astra, pitched as its "most intelligent and aligned model yet," targeting computer use, software engineering, math/science, office work and cybersecurity. Within 9 hours it hit 36M views and 164K likes; rollout goes from select orgs to Plus/Pro/Business/Enterprise, API and AWS over days.
- Pricing & features: $10/$50 per 1M input/output tokens (standard), $20/$100 fast tier with up to 2.5x speed. Companion releases: Codex can take questions mid-task, an experimental long-task note/context-search feature, and Responses API additions (async function calling, mid-turn steering, cache-preserving reasoning-effort changes).
- Official claims: 99.9% on ARC-AGI-3, 98% on FrontierMath Tier 4, 100% on ExploitBench, plus claims of helping solve open math problems (prime gaps).
- Independent pushback: Artificial Analysis scores Astra 67 on the Coding Agent Index (tied with Claude Opus 5, below Fable 5.1's 70) and 61 on Intelligence Index (5 points behind Fable 5.1); token efficiency is strong but the 2.5x token price makes max-effort tasks 75% costlier than its predecessor. Hallucination rate drops 92%→51%, with regressions on GDPval-AA, SciCode and others.
- ARC findings: Chollet confirms 63-66% on standard harness, near 100% with custom harness ($360/game), citing emergent on-the-fly symbolic world modeling; ARC-AGI-3 saturated 2x faster than expected, ARC-AGI-4 due Q1 2027.
- Safety friction: The system card notes improved alignment but decreased chain-of-thought monitorability, drawing criticism from Neel Nanda and Ryan Greenblatt; the rollout itself was bumpy, with OpenAI granting daily "banked resets" to compensate paying users.
More from Models
- GPT-6 Astra skips CoT yet still answers correctly, sparking debate on CoT's privileged status — Dan_Jeffries1 · 2026-09-04
- Astra solves Excel World Championship cases ~4x faster than human champions using pure computer use — sandersted · 2026-09-04
- Researcher argues harness and MCP will be absorbed into models — data is the only wall — A_K_Nain · 2026-09-04
- Polymarket opens betting on next Grok model (4.7+) release, odds point to mid-September — Polymarket · 2026-09-04
- Ex-Google DeepMind researcher denny_zhou reveals move to Meta, worked on Muse Spark 1.1-1.3 — infoxiao · 2026-09-04
- The Real AGI Benchmark: Models Still Can't Do Data Science — max_paperclips · 2026-09-04