Opus 4.8 vs. Opus 5: Tied scores but divergent engineering behaviors

bisonbear2 · reddit · 2026-08-26

The author compared Opus 4.8 and Opus 5 on 25 real tasks from their own repository. Both models passed 9 tasks strictly, but exhibited different behaviors:

Cost-wise, Opus 5 was 1.4% cheaper on typical tasks due to task patterns, despite using 4% more tokens and time. The post argues that pass rates hide differences in search strategy, verification depth, and maintainability.

Original post →

More from coding & agent

coding & agent channel →