Opus 5’s coding scores reportedly drop above “high” effort, not at max
hero88645 · reddit · 2026-07-25
Opus 5’s effort setting isn’t monotonic for coding
The thread argues that Opus 5 can get worse above “high” effort on coding tasks, not better. On FrontierCode, scores reportedly drop past the high setting because the model starts making unnecessary refactors and edits outside the requested scope.
It also cites two supporting data points:
- On AA-Omniscience, Opus 5 is about 11% more accurate than Opus 4.8, but its hallucination rate is roughly 6% higher.
- In CodeRabbit code review testing at xhigh, actionable-comment precision improved (39.3% vs. 35.2%), but recall on known issues fell (55.2% vs. 61.1%) and the model produced about 4× as many nitpicks.
The main takeaway is that the best effort level may be workload- and codebase-specific, and defaulting to max can waste money or even hurt output. The post also notes that when Claude’s safety classifier triggers, requests in Claude.ai / Claude Code / Cowork silently fall back to Opus 4.8 by default.
Related event: Opus 5 Coding Paradox: Higher Reasoning Leads to Lower Scores(13 posts)→
More from coding & agent
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- Steal this idea: prompt-to-hardware where agents assemble custom devices — paraschopra · 2026-09-11
- Model Is the Least Interesting Part: A Guide to Six Core AI Architectures from RAG to Multi-Agent — goyalshaliniuk · 2026-09-11
- Non-coder builds layered memory architecture: 20k tokens tracks a year of agent conversations — matteoianni · 2026-09-11