GPT-5.6 ARC-AGI-3 scores surge 3x with specific API settings enabled

sandersted · x · 2026-07-30

Testing reveals that GPT-5.6 initially struggles with the ARC-AGI-3 benchmark. However, enabling two specific API settings used internally by ChatGPT and Codex boosts its public set score by roughly 3x and increases token efficiency by 6x. This highlights that real-world performance is heavily driven by product engineering around the model, not just the base model itself.

Related event: GPT-5.6 Scores Triple on ARC-AGI-3 After Enabling Two API Settings(12 posts)→

Original post →

More from Models

Models channel →