GPT-5.6 ARC-AGI-3 Score Triples When Enabling Thought Memory

sandersted · x · 2026-07-30

An OpenAI researcher shared insights on the synergy between model capabilities and product harnesses. While GPT-5.6 performs poorly by default on the ARC-AGI-3 test, enabling two API settings used in ChatGPT and Codex (allowing the model to remember its thoughts) boosts its score by 3x and improves token efficiency by 6x.

This demonstrates that actual performance is a function of the model combined with its product harness, not just the model alone.

Related event: GPT-5.6 Scores Triple on ARC-AGI-3 After Enabling Two API Settings(12 posts)→

Original post →

More from Models

Models channel →