GPT-6 Astra Deep Dive: 99.9% ARC-AGI-3 Score Used OpenAI's Own Harness (62.7% Standard)
AGI Hunt · wechat · 2026-09-04
OpenAI's GPT-6 Astra wows with parallel seven-task demos, a first-ever DEFCON puzzle solve, and a prime-gap result (246→186) — but its 99.9% ARC-AGI-3 score used OpenAI's own harness (62.7% on the standard one), ranks only third on Artificial Analysis, and AISI found its chain-of-thought monitoring recall plummeted to 10.9%. Priced at $10/$50 per M tokens; a chaotic launch with 404s and leaked drafts.
More from Models
- System card data contradicts OpenAI's Astra alignment claim, critic says GPT-5.5 safer — GarrisonLovely · 2026-09-04
- Miles Brundage: Astra demos are crazy, Anthropic surely not far behind — Miles_Brundage · 2026-09-04
- Sean Taylor: 'Fast progress on eradicating hallucinations,' backed by realistic Astra eval — DavideCrapis · 2026-09-04
- Microsoft launches MAI-Transcribe-2, claiming 10x speed of GPT-Transcribe — ZacharyHuang12 · 2026-09-04
- Fable 5.1 likely matches Astra on CoT controllability, observers say — Miles_Brundage · 2026-09-04
- Researcher doubts Gemini outage reports: Google's in-house infra makes shared failure unlikely — generativist · 2026-09-04