ARC-AGI-3 gains on Opus 5 may reflect harness improvements more than direct targeting

mhmazur · x · 2026-07-28

Mike Knoop argues that ARC-AGI-3 was probably not directly targeted by Opus 5. Instead, he thinks the gains are coming from the model learning the Claude Code / Work harness style of operation.

His key claim is that the raw model appears sharper at strategy and multi-turn context carry-forward, which suggests the underlying API model improved, but the much larger visible jump for Claude Code/Work users may be partly an artifact of the harness already doing a lot of the heavy lifting.

He also says his earlier prediction was that most ARC v3 progress this year would come from harness improvements, and that broader on-the-fly world modeling advancements would eventually get trained directly into models, similar to how reasoning improved over time.

Original post →

More from Models

Models channel →