OpenAI: Optimized API Settings Triple GPT-5.6 Sol's Score on ARC-AGI-3

OpenAI · x · 2026-07-30

OpenAI officially revealed the latest performance of its GPT-5.6 Sol model on the ARC-AGI-3 benchmark. The research found that the model previously struggled with 2D puzzle games because the testing harness did not allow it to retain learned memory.

By enabling two API settings—retained reasoning and context compaction—the model's score jumped from 13.3% to 38.3% while using 6x fewer output tokens. This demonstrates that a better harness and context management can massively unlock a model's reasoning potential.

Related event: GPT-5.6 Scores Triple on ARC-AGI-3 After Enabling Two API Settings(17 posts)→

Original post →

More from Models

Models channel →