GPT-6 Astra's 62.7% vs 99.95% ARC-AGI-3 gap comes down to harness state persistence
HelenTKL · reddit · 2026-09-04
Two wildly different verified GPT-6 Astra scores on ARC-AGI-3 — 62.71% (Standard harness, max reasoning) and 99.95% (Provider Adapter, high reasoning) — come down to harness configuration: the Standard setup only lets the model carry forward chosen notes, while the Provider Adapter preserves opaque reasoning state between requests and uses compaction.
The poster argues: (1) the 99.95% run isn't apples-to-apples with other models' Standard scores; (2) dismissing it as 'harness cheating' misses that real agent workflows are stateful, so the model-and-harness combination may be the production capability that matters; (3) reasoning levels differ between the best runs, so it's not perfectly isolated either.
Takeaway: 'which model is smarter' is becoming an incomplete question — measured performance increasingly depends on state persistence, compaction, and the surrounding agent scaffold.
More from coding & agent
- Dev: Codex optimized for hours against the wrong label column, 'deeply uncurious' — ivan_bezdomny · 2026-09-04
- No-code dev shares full AI pipeline for cozy game: Opus, Gemini, Meshy, Claude, Unity — Gambo7592 · 2026-09-04
- 43 Projects From Singapore's ChatGPT Sites Hackathon Open for Voting — gabrielchua · 2026-09-04
- img2threejs explained: turning one photo into editable Three.js code with 80k-180k tokens — maier_ak · 2026-09-04
- Dev Adds Voice Chat to 'Ghost' AI Assistant So It Can Check His Agent Orchestrator — BLUECOW009 · 2026-09-04
- Practitioner: Tool-Call Trajectories Are Higher-Signal Than CoT for Agent Monitoring — joshua_saxe · 2026-09-04