OpenAI Reports GPT-5.6 Sol Score Triples on ARC-AGI-3 After Harness Change
daniel_mac8 · x · 2026-08-15
X user danielmac8 posts that OpenAI reported GPT-5.6 Sol jumped from 13.3% to 38.3% on ARC-AGI-3 after enabling retained reasoning and compaction, only changing the harness. This illustrates that model capability and agent capability are different, emphasizing harness engineering. Recommends following DarwinX and AutoDesign.
More from AGI Musings
- Johnny Test predicted agentic AI hacking capabilities — Casq-qsaC_178_GAP073 · 2026-08-15
- Ethan Mollick: AI product strategy predictions suffer from linear bias — eldonredwards · 2026-08-15
- Anthropic Research: Patterns and Problems in Multiagent Systems — rseroter · 2026-08-15
- Is recording everything for personal AI agents obsessive? — lucienbaba · 2026-08-15
- Anima Anandkumar: AI Models Lack Understanding of the Physical World — AnimaAnandkumar · 2026-08-15
- Former ARIA Director Davidad on Posthuman Futures: Wisdom and Cosmic Expansion — danfaggella · 2026-08-15