Same model, 54.8% to 99.9%: how OpenAI's harness sent Astra soaring on ARC-AGI-3

sourdub · reddit · 2026-09-30

A Reddit technical discussion argues that a harness doesn't make a model smarter — it just stops it from repeatedly becoming stupider. The harness is a deterministic substrate: it doesn't boost innate reasoning, but it changes the model's trajectory.

The example: ARC's standard harness lets the model keep visible notes, with the model deciding what to preserve. OpenAI's Provider Adapter for Astra, by contrast, retains reasoning state across requests and replaces rolling truncation with compaction so useful history survives long conversations. Those two changes alone took Astra from 54.8% to 99.9% on ARC-AGI-3.

Takeaway: identical weights with different state management and context strategies produce wildly different scores — harness engineering is a major benchmark variable.

Original post →

More from coding & agent

coding & agent channel →