CodeRabbit says Opus 5 misses more bugs even as it costs more tokens per call
socialwithaayan · x · 2026-07-27
A long thread argues that model evals are only useful if you add a verification layer.
- It says CodeRabbit tested Opus 5 on real pull requests and found it makes fewer wrong claims than GPT-5.6 Sol, but misses more bugs.
- The thread also claims Opus 5 reads about 50% more and writes about 65% more tokens per call than baseline models, so the cost premium is in volume rather than just the rate card.
- The proposed workflow is to break work into isolated claims, challenge each claim with the strongest objection, and then source-check the result before shipping.
- It closes with a content-pipeline prompt that tells a model to learn from five top posts, extract the pattern, and then generate in that style.
More from coding & agent
- Moonshot open-sources AgentENV, a distributed platform for large-scale agent training — teortaxesTex · 2026-07-27
- New study says agent skills should be judged by regressions, not just average gains — omarsar0 · 2026-07-27
- A Claude Opus 5 prompt reportedly built a playable birthday game in a 21-hour run — mattshumer_ · 2026-07-27
- A new skill cuts Claude.md clutter by 50% and audits agent instructions — iamrobotbear · 2026-07-27
- Anthropic ships a beta security plugin for Claude Code with multi-agent scans — thione · 2026-07-27
- Poolside releases Laguna S 2.1, an open-weight coding model with 1M context — thione · 2026-07-27