ARC-AGI-3 gains on Opus 5 may reflect harness improvements more than direct targeting
mhmazur · x · 2026-07-28
Mike Knoop argues that ARC-AGI-3 was probably not directly targeted by Opus 5. Instead, he thinks the gains are coming from the model learning the Claude Code / Work harness style of operation.
His key claim is that the raw model appears sharper at strategy and multi-turn context carry-forward, which suggests the underlying API model improved, but the much larger visible jump for Claude Code/Work users may be partly an artifact of the harness already doing a lot of the heavy lifting.
He also says his earlier prediction was that most ARC v3 progress this year would come from harness improvements, and that broader on-the-fly world modeling advancements would eventually get trained directly into models, similar to how reasoning improved over time.
More from Models
- Kimi report reveals a wide internal benchmark suite for coding and agent skills — stochasticchasm · 2026-07-28
- Claude is still being called the most steerable model set, despite its weirdness — sloppenheimer · 2026-07-28
- Frontend Code Arena: Opus 5 Max Takes #1, Kimi K3 Max Follows Closely — arena · 2026-07-28
- Kimi K3 Max Tops Arena Leaderboard in Frontend Code and Agent Tasks — arena · 2026-07-28
- Kimi K3 Hits Hugging Face Inference API at $15/M Output Tokens — mervenoyann · 2026-07-28
- OpenRouter cuts GPT-5.6 Terra and Luna prices by 50% — OpenAIDevs · 2026-07-28