Astra's 98.6% ARC-AGI-3 score called misleading: fair comparison shows just 54.8%
PsychologicalSoup251 · reddit · 2026-09-04
A Reddit analysis argues OpenAI's reported 98.6% ARC-AGI-3 score for Astra is technically true but deliberately misleading:
- Uneven conditions: Astra used a custom agentic harness at MAX thinking, while GPT 5.6 Sol (7.8%) and Claude Opus 5 (30.2%) ran standard harnesses, with Opus 5 only at HIGH thinking.
- Saturation: Under the same custom-harness rules, older GPT 5.6 Sol has already scored 99.0–99.9%, and one harness pushed both Sol and Opus 5 to 100.0%.
- Fair comparison: On the standard harness at HIGH thinking, Astra scores 54.8% vs Opus 5's 30.2% — a real but far smaller lead.
The author concludes ARC-AGI-3 is oversaturated and fundamentally harness-dependent, useful only for measuring LLM+harness combos rather than raw model ability.
More from Models
- System card data contradicts OpenAI's Astra alignment claim, critic says GPT-5.5 safer — GarrisonLovely · 2026-09-04
- Miles Brundage: Astra demos are crazy, Anthropic surely not far behind — Miles_Brundage · 2026-09-04
- Microsoft launches MAI-Transcribe-2, claiming 10x speed of GPT-Transcribe — ZacharyHuang12 · 2026-09-04
- Fable 5.1 likely matches Astra on CoT controllability, observers say — Miles_Brundage · 2026-09-04
- Researcher doubts Gemini outage reports: Google's in-house infra makes shared failure unlikely — generativist · 2026-09-04
- ThursdAI breaks down OpenAI GPT-6 Astra: 99% on Arc-AGI, standout computer use — altryne · 2026-09-04