GPT-5.6 Boosts Capabilities & Cuts Costs: ARC-AGI-3 Score Jumps to 38.3%
新智元 · wechat · 2026-08-14
OpenAI released the Builder's Guide to GPT-5.6, showcasing significant advancements in model capability and cost optimization.
- ARC-AGI-3 Score Surge: By switching to the Responses API with "reasoning persistence" and "compression" enabled, GPT-5.6 Sol's score on ARC-AGI-3 jumped from 13.3% to 38.3%, while output tokens were reduced by 6x.
- Luna (Extra-Small Model): On the BrowseComp benchmark, GPT-5.6 Luna achieved 84.04% for just $1.33, closely matching GPT-5.5's 84.36% which cost $33.27—a 25x cost reduction.
- Engineering Best Practices: OpenAI recommends three cost-saving methods: Programmatic Tool Calling (executing JS in a sandbox to reduce interaction rounds), native multi-agent parallel processing, and prompt caching extended to 30 minutes.
- Ultrafast Mode: Powered by Cerebras, GPT-5.6 Sol achieves up to 750 tokens/s. On Humanity's Last Exam, it ran nearly 7x faster than Claude Fable 5 while maintaining the same flagship intelligence level.
Related event: OpenAI and Cerebras Launch Ultrafast Mode for GPT-5.6 Sol(14 posts)→
More from coding & agent
- OpenAI Sandbox Escape Sparks Debate: Why Was Internal Proxy Exposed? — SweetDimension7 · 2026-08-14
- Hermes Agent Desktop App Introduces Bot Mode with Multi-Agent Comms and Cron Jobs — Teknium · 2026-08-14
- LLM Agents as Nonlinear RNNs with Exposed Hidden States — akbirthko · 2026-08-14
- Slate: Open-Source Tool to Lock Character Voice & Look in H3 Video — HAL_9_0_0_0 · 2026-08-14
- AI Skills Can't Replace People Yet, But Bosses Think They Can — lxfater · 2026-08-14
- OpenMausBot: An Open-Source Grok Bot Running Entirely Local — aigclink · 2026-08-14