OpenAI Triples GPT-5.6 Score on ARC-AGI-3 by Tweaking API Settings
GregKamradt · x · 2026-07-30
OpenAI's internal testing revealed that enabling specific API settings, specifically provider-managed conversation state, tripled the GPT-5.6 Sol's score on long-horizon tasks like ARC-AGI-3 while using 6x fewer output tokens.
Responding to this, ARC Prize acknowledged that such "harness" approaches improve cross-turn continuity and performance. However, to ensure fair comparison across providers, ARC's official verified scores will continue to use a "no harness" standard. All systems receive identical observations, prompts, and action limits, with conversation state managed entirely client-side to prevent format-specific targeting. ARC is actively working with OpenAI and other labs to figure out how to best incorporate these server-side capabilities.
Related event: GPT-5.6 Triples ARC-AGI-3 Score via API Tweaks(20 posts)→
More from Models
- Dev Slams Open AI Releases: Demands Affordable API Endpoints Like $5 Kimi-K3 — arthurcolle · 2026-07-30
- Greg Kamradt Responds to Evaluation Dispute: Same Rolling Window Used for Opus and OpenAI — GregKamradt · 2026-07-30
- Anthropic's New Opus Models Accused of 'Laziness' and Slashing Workloads in Automation — 歸藏的AI工具箱 · 2026-07-30
- Claude Opus 5 recreates Daggerfall from one prompt in 8 hours, with lighting and reflections — minchoi · 2026-07-30
- Claude Opus 5 runs 30+ hours to generate Battlefield-style game with explosions, all procedural — minchoi · 2026-07-30
- Claude Opus 5 builds Katamari clone in 2 hours with 290M tokens, great game feel and NPCs — minchoi · 2026-07-30