ARC Benchmark: Frontier Models Absorbing Harness Patterns for Better Reasoning

mhmazur · x · 2026-08-14

Following the launch of the ARC v3 benchmark, the community has seen rapid progress, with many harnesses on the leaderboard now scoring over 90% on the public set.

Mike Knoop, co-founder of ARC, notes that the key trend is that models are now training in these useful harness patterns. For instance, Claude 3 Opus effectively emulates on-the-fly world modeling and strategy carry-forward, achieving around 30% on the semi-private set. To preserve the benchmark's signal on whether AI is genuinely getting smarter, the official policy strictly prohibits verifying community harnesses on the private set to prevent over-targeting.

Original post →

More from Models

Models channel →