GPT-6 'Astra' Smashes ARC-AGI Records: 62.7% on Standard Harness, 99.9% on ARC-AGI-3
repligate · x · 2026-09-04
mhmazur of the ARC-AGI team published a detailed analysis of GPT-6 "Astra," calling it a career highlight:
- On the standard harness (which lets models carry notes forward, historically used for apples-to-apples comparisons), Astra scored 62.7% — more than doubling the previous verified high — and set a new record on ARC-AGI-2.
- With a new provider adapter harness that preserves the model's opaque reasoning state between requests and uses native compaction as context grows, Astra hit 99.9% on ARC-AGI-3.
- Going forward, the team will test models with both harnesses.
Commenters note Astra appears to "bend the cost/performance continuum backwards," suggesting unusually strong value for money.
Related event: GPT-6 Astra Sets ARC-AGI Records, Scoring 99.9% on ARC-AGI-3(3 posts)→
More from Models
- Critics warn OpenAI's GPT-6 Astra reasons opaquely, gutting CoT monitoring safety — GaryMarcus · 2026-09-04
- ChatGPT adds writing-style matching from connected apps, analytics, and a Yubikey deal tied to Daybreak access — btibor91 · 2026-09-04
- GPT-6 reportedly launches as Tesla starts public rides in steering-free Cybercab — Dr_Singularity · 2026-09-04
- Astra early-access users' similar blender demos look coordinated, with no practical examples shown — jdjohnson · 2026-09-04
- Researcher teases dynamic composite eval index as "evals run on Twitter vibes" — evijit · 2026-09-04
- OpenAI engineer says GPT-6 "Astra" has "tremendously" improved writing quality — mark_k · 2026-09-04