CodeRabbit says Opus 5 misses more bugs even as it costs more tokens per call
socialwithaayan · x · 2026-07-27
A long thread argues that model evals are only useful if you add a verification layer.
- It says CodeRabbit tested Opus 5 on real pull requests and found it makes fewer wrong claims than GPT-5.6 Sol, but misses more bugs.
- The thread also claims Opus 5 reads about 50% more and writes about 65% more tokens per call than baseline models, so the cost premium is in volume rather than just the rate card.
- The proposed workflow is to break work into isolated claims, challenge each claim with the strongest objection, and then source-check the result before shipping.
- It closes with a content-pipeline prompt that tells a model to learn from five top posts, extract the pattern, and then generate in that style.
Related event: Opus 5 Code Audit: 1M Context Comes at Higher Token Cost(2 posts)→
More from coding & agent
- Dev claims 20k more commits coming: Opus 5.5 and GPT-6 Sol supercharge his output — doodlestein · 2026-09-23
- A JEV-powered Wireshark classifier accidentally uncovered real backdoors on a home network — multiply_matrix · 2026-09-23
- 299 real intents tested: classifier routing trails GLM-4-Flash by 3 points but is 6.5x faster — Sufficient_Flower860 · 2026-09-23
- OpenExecutive: open-source virtual executive team of 8 specialist AI agents hits 5.1k GitHub stars — tom_doerr · 2026-09-23
- Framer launches Skills: teach your design agent reusable workflows, design systems and CMS rules — soleio · 2026-09-23
- Cursor, OpenAI and Anthropic shipped coordinator-agent fleets in one week, but the review bottleneck stays — omidfarhang · 2026-09-23