FrontierCode 1.1 says Opus 5 drops in score at xhigh and max reasoning settings
inductionheads · x · 2026-07-25
FrontierCode says higher reasoning hurts Opus 5 on some criteria
A quoted note says Opus 5 is strong and roughly delivers Fable-level performance at Opus pricing, but that at higher reasoning effort it regresses on some of the criteria used in FrontierCode.
The linked investigation says Opus 5 actually scores lower on FrontierCode 1.1 at xhigh and max reasoning settings. The key issue appears to be the benchmark’s scope metric, which penalized the model for making edits unrelated to the task.
So the headline result is not just about correctness: the evaluation suggests that when reasoning is pushed harder, the model can become less aligned with the benchmark’s broader task discipline.
Related event: Opus 5 Coding Paradox: Higher Reasoning Leads to Lower Scores(13 posts)→
More from Models
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11