OpenAI Triples GPT-5.6 Score on ARC-AGI-3 by Tweaking API Settings

GregKamradt · x · 2026-07-30

OpenAI's internal testing revealed that enabling specific API settings, specifically provider-managed conversation state, tripled the GPT-5.6 Sol's score on long-horizon tasks like ARC-AGI-3 while using 6x fewer output tokens.

Responding to this, ARC Prize acknowledged that such "harness" approaches improve cross-turn continuity and performance. However, to ensure fair comparison across providers, ARC's official verified scores will continue to use a "no harness" standard. All systems receive identical observations, prompts, and action limits, with conversation state managed entirely client-side to prevent format-specific targeting. ARC is actively working with OpenAI and other labs to figure out how to best incorporate these server-side capabilities.

Related event: GPT-5.6 Triples ARC-AGI-3 Score via API Tweaks(20 posts)→

Original post →

More from Models

Models channel →