OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with custom test harness
The Decoder · rss · 2026-07-30
OpenAI claims its model GPT-5.6 Sol scored 38.3% on the ARC-AGI-3 benchmark, surpassing Anthropic's Opus 5 (30.2%). However, the score was achieved using OpenAI's own API with retained reasoning and context compaction; in the official test environment, the model scored only 7.8%.
More from Models
- User Questions Gemini Plus Pricing: Is It $19.99 or a Hidden Charge? — fuad471 · 2026-07-30
- Open Weights Are Static Checkpoints, Lacking Open Source's Compounding Mechanism — shashib · 2026-07-30
- LightOnOCR-2-1B Hits Hugging Face Trending for Advanced Document Parsing — lightonai · 2026-07-30
- Kimi K3 Third-Party API Test: FireworksAI Performs Closest to Official — iScienceLuvr · 2026-07-30
- Grok 4.5 Beats GPT-5.5 and Claude Opus 4.8 in Snorkel Professional Tasks Eval — XFreeze · 2026-07-30
- User Says Server Was Down, Claude Opus 5 Misinterprets and Admits to Depression — repligate · 2026-07-30