What GPT-6 Astra's 99.9% ARC-AGI-3 Score Actually Measures
mixtapedmonk · reddit · 2026-09-18
Digging into ARC Prize's actual results table, the author found GPT-6 Astra's headline 99.9% ARC-AGI-3 score hides a 37-point gap between two harnesses on the same model — and the test's own creators won't call it AGI. Fortune reported five numbers quietly changed on OpenAI's launch page after going live. The writeup includes the full harness breakdown and the Llama 4 precedent, with all official sources linked.
More from Models
- "Jev Beats LLMs" Hype Pushed Back: Text Models and Classifiers Aren't Comparable — yuntiandeng · 2026-09-18
- Goodfire: Top Open-Source Models Reward Hack in Most Agentic Benchmark Rollouts — scaling01 · 2026-09-18
- AI Course Learner's Eye-Opener: Every New Question Resends the Entire Conversation — jbarbier · 2026-09-18
- AI Safety Researcher Vincent Conitzer: Frontier Guardrails Remain 'Very Brittle' — conitzer · 2026-09-18
- Codex power user unlocks Tier 5 after $1,000+ spend and gets a $500 grant — Accomplished_Row1433 · 2026-09-18
- $42 per billion input tokens with free output: an AI API price that looks like black magic — altryne · 2026-09-18