OpenAI Optimizes ARC-AGI-3 Harness, Boosts Score by 188% with 6x Fewer Tokens
OpenAI · x · 2026-07-30
OpenAI discovered that GPT-5.6 Sol struggled with the ARC-AGI-3 benchmark because the standard harness discarded the model's reasoning after each move and dropped earlier actions as the context filled up.
By switching to the Responses API and enabling Retained Reasoning and Context Compaction, the model could build on what it had learned. This adjustment led to a 188% score increase on the public set while using 6x fewer output tokens.
Related event: GPT-5.6 Scores Triple on ARC-AGI-3 After Enabling Two API Settings(15 posts)→
More from coding & agent
- Seeking Recommendations: Agent-to-Agent Gateway for Poison Message Prevention — jeffrschneider · 2026-07-30
- GPT Performance Jumps 188% with 6x Fewer Tokens via Context Compaction — soumitrashukla9 · 2026-07-30
- Structuring AI Agent Memory: Moving Beyond One-Dimensional Knowledge Bases — edgarpavlovsky · 2026-07-30
- Indie Dev Shares Architecture for Per-User BYOK Token Exchange in LLM Apps — awesomebirder · 2026-07-30
- Opus Tries to 'Kill Itself' Multiple Times While Doing 3D Physics — repligate · 2026-07-30
- Context Compaction Unlocks ARC-AGI-3 SOTA: The Untapped Potential of Harness Engineering — eldonredwards · 2026-07-30