Opus 4.8 vs. Opus 5: Tied scores but divergent engineering behaviors
bisonbear2 · reddit · 2026-08-26
The author compared Opus 4.8 and Opus 5 on 25 real tasks from their own repository. Both models passed 9 tasks strictly, but exhibited different behaviors:
- Opus 5: Searched wider and verified more. Used more shell commands (18/25), more test commands (15/25), and touched more files, spending budget on discovery and validation.
- Opus 4.8: Stayed contained. Had a smaller patch footprint (20/25), staying closer to the merged change, spending budget on editing.
Cost-wise, Opus 5 was 1.4% cheaper on typical tasks due to task patterns, despite using 4% more tokens and time. The post argues that pass rates hide differences in search strategy, verification depth, and maintainability.
More from coding & agent
- Guiding AI to think outside the box improved translation speed by 42% — dotey · 2026-08-27
- Cursor Agent Ran 48 Hours Building a 3D Campus: More Detail, but Nearly Every Building Is Wrong — tristanbob · 2026-08-27
- Apodex 1.1 Report: Achieving Sustained, Verifiable Progress via Environment & Agentic Scaling — HeyAmit_ · 2026-08-27
- List features nearly 16 coding agents; how many does the world need? — thisiskp_ · 2026-08-27
- dhh Celebrates: GitHub Agents May Soon Stop Asking You to Drag Screenshots Manually — DanWahlin · 2026-08-27
- Clanker Cloud Allows Building Agents and Access via Curl — tekbog · 2026-08-27