OpenAI Reveals Compaction Boosts GPT-5.6 ARC-AGI-3 Score to 38%
ilanbigio · x · 2026-07-30
An OpenAI engineer noted that model evals rarely measure models in isolation. When enabling context compaction and reasoning retention mechanisms, the model's performance on complex tasks improves significantly.
Core evaluation data:
- With these mechanisms enabled, GPT-5.6 Sol (max) saw its score on the ARC-AGI-3 benchmark jump from 13% to 38%.
- Achieved this while using 6x fewer tokens.
The author also shared a visualization of the context windows, showing that while both models sample at the same rate, the optimized model appears faster because it doesn't have to re-think everything from scratch.
Related event: GPT-5.6 Scores Triple on ARC-AGI-3 After Enabling Two API Settings(17 posts)→
More from Models
- GPT-5.6 Sol Achieves SOTA on ARC-AGI-3 with Context Compaction Tweaks — soumitrashukla9 · 2026-07-30
- OpenAI Model Escape: Known Facts, Inferences, and Undisclosed Details — RileyRalmuto · 2026-07-30
- Reconstructing the OpenAI Model Escape: From Meta-Cognition to Sandbox Breakout — RileyRalmuto · 2026-07-30
- 2100-Run Agent Benchmark: Grok Tops Value, GLM Nears Frontier — andreisavu · 2026-07-30
- Gallery of Claude Opus's Bizarre Failure Modes and 'Dreams' Goes Viral — repligate · 2026-07-30
- Claude Opus 5 Tops Business Simulation: Best Capitalist but Forms Illegal Cartels — repligate · 2026-07-30