GPT-5.6 Scores Triple on ARC-AGI-3 After Enabling Two API Settings
GPT-5.6's performance on the ARC-AGI-3 benchmark varies drastically depending on API configurations. Developer tests reveal that enabling two specific settings used internally by ChatGPT and Codex boosts the model's score by roughly 3x to 38%, while improving token efficiency by about 6x. This highlights the critical synergy between model capabilities and product frameworks, suggesting that isolated evaluations may severely underestimate a model's true potential.
已确认
- Developer @sandersted found that GPT-5.6 performs poorly by default on the ARC-AGI-3 benchmark
- Enabling two specific API settings used internally by ChatGPT and Codex increased the public test set score by about 3x, reaching 38%
- Token efficiency simultaneously improved by roughly 6x
- These two settings are 上下文压缩(compaction) and 推理记忆保留(reasoning retention), allowing the model to remember previous thoughts
- OpenAI engineer @ilanbigio pointed out that model evaluations should generally not be conducted in isolated environments
为什么重要
- This discovery challenges the industry's current isolated evaluation methods, suggesting they may fail to reflect a model's actual capabilities in real-world product environments
- The deep synergy between product frameworks (such as ChatGPT and Codex's internal mechanisms) and model capabilities has a decisive impact on performance in complex tasks
- This also implies that simply comparing scores from bare-bones benchmark runs may not accurately measure the true intelligence of different AI systems
2026-07-30 ~ 2026-07-30 · 14 related posts
- Episode 1: OpenAI Says GPT-5.6 Sol Self-Optimizes, Cutting Inference Costs by 20%(2026-07-30, 8 posts)
- Episode 2: GPT-5.6 Scores Triple on ARC-AGI-3 After Enabling Two API Settings(2026-07-30, 14 posts)
Primary sources
- GPT-5.6 ARC-AGI-3 scores surge 3x with specific API settings enabled — sandersted · 2026-07-30
- GPT-5.6 ARC-AGI-3 Score Jumps 3x With Specific API Settings — sandersted · 2026-07-30
- GPT-5.6 ARC-AGI-3 Score Triples When Enabling Thought Memory — sandersted · 2026-07-30
- [source] OpenAI Reveals Compaction Boosts GPT-5.6 ARC-AGI-3 Score to 38% — ilanbigio · 2026-07-30
- Optimized Harness Achieves SoTA on ARC-AGI-3, Questioning Benchmark Validity — Angaisb_ · 2026-07-30
- OpenAI Tripled Its ARC-AGI-3 Scores by Enabling Just Two Settings — ObiWanCanownme · 2026-07-30
- [source] OpenAI: Optimized API Settings Triple GPT-5.6 Sol's Score on ARC-AGI-3 — OpenAI · 2026-07-30
- GPT-5.6 Sol Achieves SOTA on ARC-AGI-3 via Context Compaction — daniel_mac8 · 2026-07-30
- Retaining Reasoning Across Turns Boosts GPT-5.6 Sol on ARC-AGI-3 by ~3x — charliermarsh · 2026-07-30
- Harness Matters: GPT-5.6 Hits SoTA on ARC-AGI-3 with Two Setting Tweaks — TheZachMueller · 2026-07-30
4 near-duplicate retellings: sandersted · ilanbigio · OpenAI · tw_killian