OpenAI Triples ARC-AGI-3 Score with 6x Fewer Tokens via API Tweaks
tw_killian · x · 2026-07-30
OpenAI explained why GPT-5.6 Sol, despite solving open math problems, struggled with the ARC-AGI-3 benchmark of 2D puzzle games. They found the harness was not letting the model remember what it had learned. By enabling two API settings, the score tripled with 6x fewer output tokens. Researchers subsequently referenced related work showing that using exactly these two levers in text game environments has yielded strong results.
Related event: GPT-5.6 Scores Triple on ARC-AGI-3 After Enabling Two API Settings(18 posts)→
More from coding & agent
- Clarifying Agent Concepts: Differences Between MCP, Skill, and Tool Call — yangyi · 2026-07-30
- Workflow Share: Using Opus for Planning and Grok Subagents for Coding — kevinnbass · 2026-07-30
- Mnemos Engine: Building a Continuity and Accountability Framework for Autonomous Agents — RileyRalmuto · 2026-07-30
- 2100-Run Agent Benchmark: Grok Tops Value, GLM Nears Frontier — andreisavu · 2026-07-30
- Nadella Demos New Copilot Skills: Building Enterprise ROIC App with One Prompt — satyanadella · 2026-07-30
- Seeking Recommendations: Agent-to-Agent Gateway for Poison Message Prevention — jeffrschneider · 2026-07-30