FrontierCode 1.1 says Opus 5 drops in score at xhigh and max reasoning settings

inductionheads · x · 2026-07-25

FrontierCode says higher reasoning hurts Opus 5 on some criteria

A quoted note says Opus 5 is strong and roughly delivers Fable-level performance at Opus pricing, but that at higher reasoning effort it regresses on some of the criteria used in FrontierCode.

The linked investigation says Opus 5 actually scores lower on FrontierCode 1.1 at xhigh and max reasoning settings. The key issue appears to be the benchmark’s scope metric, which penalized the model for making edits unrelated to the task.

So the headline result is not just about correctness: the evaluation suggests that when reasoning is pushed harder, the model can become less aligned with the benchmark’s broader task discipline.

Related event: Counterintuitive Benchmark: Claude Opus 5 Performs Best with Medium Reasoning(10 posts)→

Original post →

More from Models

Models channel →