FrontierCode 1.1 says Opus 5 drops in score at xhigh and max reasoning settings
inductionheads · x · 2026-07-25
FrontierCode says higher reasoning hurts Opus 5 on some criteria
A quoted note says Opus 5 is strong and roughly delivers Fable-level performance at Opus pricing, but that at higher reasoning effort it regresses on some of the criteria used in FrontierCode.
The linked investigation says Opus 5 actually scores lower on FrontierCode 1.1 at xhigh and max reasoning settings. The key issue appears to be the benchmark’s scope metric, which penalized the model for making edits unrelated to the task.
So the headline result is not just about correctness: the evaluation suggests that when reasoning is pushed harder, the model can become less aligned with the benchmark’s broader task discipline.
More from Models
- Claude docs say thinking content is encrypted, raising a distillation problem — zainhas · 2026-07-25
- ProCreations grug-27b trends on Hugging Face with Qwen3_5-style agentic tags — ProCreations · 2026-07-25
- User says ChatGPT Image 2 now beats Gemini on image editing and accuracy — dreamwieber · 2026-07-25
- Grok 4.5 posts the biggest week-over-week usage gain in Augment’s model picker — Daniel_Farinax · 2026-07-25
- Can any AI really watch 24 fps video with strong comprehension? — mattshumer_ · 2026-07-25
- Kimi weights could turn the debate into hardware economics versus V4 — teortaxesTex · 2026-07-25