GPT-6 Astra's 62.7% vs 99.95% ARC-AGI-3 gap comes down to harness state persistence

HelenTKL · reddit · 2026-09-04

Two wildly different verified GPT-6 Astra scores on ARC-AGI-3 — 62.71% (Standard harness, max reasoning) and 99.95% (Provider Adapter, high reasoning) — come down to harness configuration: the Standard setup only lets the model carry forward chosen notes, while the Provider Adapter preserves opaque reasoning state between requests and uses compaction.

The poster argues: (1) the 99.95% run isn't apples-to-apples with other models' Standard scores; (2) dismissing it as 'harness cheating' misses that real agent workflows are stateful, so the model-and-harness combination may be the production capability that matters; (3) reasoning levels differ between the best runs, so it's not perfectly isolated either.

Takeaway: 'which model is smarter' is becoming an incomplete question — measured performance increasingly depends on state persistence, compaction, and the surrounding agent scaffold.

Related event: GPT-6 Astra Launch: Capability Leap Marred by Benchmarking Dispute and Declining Monitorability(62 posts)→

Original post →

More from coding & agent

coding & agent channel →