Claude Opus 5 scores better on FrontierCode at medium reasoning than at max

zainhas · x · 2026-07-25

A user asks whether Claude Opus 5 has an overthinking problem. The benchmark snippet suggests that on FrontierCode main and extended sets, Opus 5 at medium reasoning effort performs better than at high, xhigh, and max.

The implication is that pushing reasoning effort higher may hurt score on this benchmark, raising the familiar question of whether more deliberation always helps—or can sometimes become pay-to-lose.

Related event: Claude Opus 5 Peaks at Medium Reasoning in FrontierCode(5 posts)→

Original post →

More from Models

Models channel →