OpenAI Triples ARC-AGI-3 Score with 6x Fewer Tokens via API Tweaks

tw_killian · x · 2026-07-30

OpenAI explained why GPT-5.6 Sol, despite solving open math problems, struggled with the ARC-AGI-3 benchmark of 2D puzzle games. They found the harness was not letting the model remember what it had learned. By enabling two API settings, the score tripled with 6x fewer output tokens. Researchers subsequently referenced related work showing that using exactly these two levers in text game environments has yielded strong results.

Related event: GPT-5.6 Scores Triple on ARC-AGI-3 After Enabling Two API Settings(18 posts)→

Original post →

More from coding & agent

coding & agent channel →