GPT-5.6 Sol improves ARC-AGI-3 by 25% with 6x fewer tokens
OpenAIDevs · x · 2026-08-17
GPT-5.6 Sol improved its accuracy on ARC-AGI-3 from 13.3% to 38.3% using retained reasoning and compaction, while consuming roughly 6x fewer output tokens. Additionally, real estate startup Hypha AI maintained 98% of GPT-5.5's accuracy at 1/18 the cost with GPT-5.6 Luna for document extraction, and Rogo AI used programmatic tool calling to match evaluation quality with 21% fewer input tokens.
Related event: OpenAI Details GPT-5.6 Efficiency and Cost Gains(3 posts)→
More from Models
- Grok 4.6 coding test: nails complex logic but systematically misses basics — mark_k · 2026-08-18
- Qwen3.8-27B Benchmarks Show It Neck and Neck with DeepSeek V4 and GPT-5.6 Luna Max — anderspitman · 2026-08-18
- Orion 16B hits 100B training tokens using DPP on distributed GPUs — markjeffrey · 2026-08-18
- RareBench results: Gemini 3.7 Flash leaps ahead, DeepSeek Pro shows no gain — danielmckinn0n · 2026-08-18
- DeepSeek V4-Pro goes GA with configurable reasoning effort and Responses API — thione · 2026-08-18
- xAI releases Grok 4.6, focused on long-running agents, matching GPT-5.6 Sol on AA index — thione · 2026-08-18