Same model, 54.8% to 99.9%: how OpenAI's harness sent Astra soaring on ARC-AGI-3
sourdub · reddit · 2026-09-30
A Reddit technical discussion argues that a harness doesn't make a model smarter — it just stops it from repeatedly becoming stupider. The harness is a deterministic substrate: it doesn't boost innate reasoning, but it changes the model's trajectory.
The example: ARC's standard harness lets the model keep visible notes, with the model deciding what to preserve. OpenAI's Provider Adapter for Astra, by contrast, retains reasoning state across requests and replaces rolling truncation with compaction so useful history survives long conversations. Those two changes alone took Astra from 54.8% to 99.9% on ARC-AGI-3.
Takeaway: identical weights with different state management and context strategies produce wildly different scores — harness engineering is a major benchmark variable.
More from coding & agent
- Denying GPT-6.1 Sol write access as orchestrator cut coding costs 77%, at 6x runtime — GapNew4766 · 2026-09-30
- 8 research agents self-train a 30B model for 144 hours in RSIArena livestream experiment — my_cat_can_code · 2026-09-30
- OpenAI's Nan Yu: a boss agent running other agents is just one agent with extra steps — victor_explore · 2026-09-30
- Open-source MCP server lets agents query 124.6B TikTok data points — operatorarkay · 2026-09-30
- Dev argues truly always-on autonomous agents have never actually been tried — jacob_posel · 2026-09-30
- Skill scaffold template: same shape, faster review, fewer broken CIs — blaizedsouza · 2026-09-30