Opus 5 beats higher effort on FrontierCode at medium effort

kieranklaassen · x · 2026-07-25

A repost of a thread claiming that Opus 5 scores better on FrontierCode at medium effort than at higher effort, even though effort still improves results on some other evals.

The implication is that more compute or longer deliberation does not always translate into better coding performance on every benchmark, and FrontierCode may be exposing a different failure mode than other tests.

Related event: Counterintuitive Benchmark: Claude Opus 5 Performs Best with Medium Reasoning(10 posts)→

Original post →

More from Models

Models channel →