Claude Opus 5 scores better on FrontierCode at medium reasoning than at max
zainhas · x · 2026-07-25
A user asks whether Claude Opus 5 has an overthinking problem. The benchmark snippet suggests that on FrontierCode main and extended sets, Opus 5 at medium reasoning effort performs better than at high, xhigh, and max.
The implication is that pushing reasoning effort higher may hurt score on this benchmark, raising the familiar question of whether more deliberation always helps—or can sometimes become pay-to-lose.
Related event: Claude Opus 5 Peaks at Medium Reasoning in FrontierCode(5 posts)→
More from Models
- Claude Opus 5 reportedly beats Fable 5 on a hard 3D coding test at 75% of the price — rohanpaul_ai · 2026-07-25
- Frontier lab rumor says Opus 5 ARC-AGI 3 score looks fake — flowersslop · 2026-07-25
- Chart says Claude Opus 5 blocks far less defensive coding than Fable 5 — repligate · 2026-07-25
- Claude Opus 5 lands, with DirectTerminal bringing richer Claude Code output to the terminal — draginol · 2026-07-25
- A Codex reset calendar shows usage limits do not always refresh at midnight UTC — petrusenko_max · 2026-07-25
- Grok 4.5 tops Ramp’s invoice test on 150,000 real business bills — elonmusk · 2026-07-25