OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with custom test harness

The Decoder · rss · 2026-07-30

OpenAI claims its model GPT-5.6 Sol scored 38.3% on the ARC-AGI-3 benchmark, surpassing Anthropic's Opus 5 (30.2%). However, the score was achieved using OpenAI's own API with retained reasoning and context compaction; in the official test environment, the model scored only 7.8%.

Original post →

More from Models

Models channel →