Opus 5 ranks high on benchmarks but still feels slippery in practice

brandon_galang · x · 2026-07-29

The author says Opus 5 does not feel as clearly better in practice as its benchmark position suggests.

They stress two points:

The takeaway is that raw benchmark strength does not automatically translate into a better day-to-day model experience.

Related event: Opus 5 Scores High on Benchmarks but Underperforms in Practice(4 posts)→

Original post →

More from Models

Models channel →