Opus 5’s coding scores reportedly drop above “high” effort, not at max
hero88645 · reddit · 2026-07-25
Opus 5’s effort setting isn’t monotonic for coding
The thread argues that Opus 5 can get worse above “high” effort on coding tasks, not better. On FrontierCode, scores reportedly drop past the high setting because the model starts making unnecessary refactors and edits outside the requested scope.
It also cites two supporting data points:
- On AA-Omniscience, Opus 5 is about 11% more accurate than Opus 4.8, but its hallucination rate is roughly 6% higher.
- In CodeRabbit code review testing at xhigh, actionable-comment precision improved (39.3% vs. 35.2%), but recall on known issues fell (55.2% vs. 61.1%) and the model produced about 4× as many nitpicks.
The main takeaway is that the best effort level may be workload- and codebase-specific, and defaulting to max can waste money or even hurt output. The post also notes that when Claude’s safety classifier triggers, requests in Claude.ai / Claude Code / Cowork silently fall back to Opus 4.8 by default.
Related event: Opus 5 Coding Scores Drop with Higher Reasoning Effort(11 posts)→
More from coding & agent
- 6 weeks, 100 lessons: how to build a personal AI agent that actually remembers and plans — Div_pradeep · 2026-07-25
- Local agent beats Hermes on GAIA Level 1 while running fully in llama.cpp — HeyAmit_ · 2026-07-25
- Every Codex or Claude Code complaint turns into a product pitch in the replies — dejavucoder · 2026-07-25
- Grok teases a modular VS Code extensions model for extensible agent fleets — dee_hw · 2026-07-25
- Animam ships a multi-tenant AI agent platform with widget, API, voice, and MCP — animam-tech · 2026-07-25
- Andrew Chen asks whether anyone is actually coding inside Claude and ChatGPT desktop apps — andrewchen · 2026-07-25