GPT-6 Astra maxes ARC-AGI-3 with compaction, but harnesses differ
mhmazur · x · 2026-09-04
Steven Heidel (OpenAI) showed GPT-6 Astra maxing out ARC-AGI-3 when compaction is enabled in the Responses API, beating Opus 5 and GPT-5.6 Sol. But mhmazur notes the comparison isn't apples to apples: Astra's score uses the new Provider Adapter harness (preserving opaque reasoning across requests and compaction), while Opus 5 and GPT-5.6 Sol were run on the Standard harness.
More from Models
- OpenAI launches GPT-6 Astra, claiming it can do anything you do on a computer — kagigz · 2026-09-04
- Databricks evals: GPT-6 Astra claims SOTA on OfficeQA Pro benchmarks, cheaper per task — downingARK · 2026-09-04
- Mathematician tests GPT-6 Astra: live Lean proof verification while writing arguments — teortaxesTex · 2026-09-04
- Tavus Launches Sparrow-2, Claiming #1 in End-of-Turn Detection and Interruption Handling — ycombinator · 2026-09-04
- ARC-AGI-3: 10x reasoning tokens cuts total cost from $48k to $26k vs medium — i_dg23 · 2026-09-04
- Researchers flag data contamination concerns in benchmark behind Astra's time-horizon score — dfrsrchtwts · 2026-09-04