Opus 5 max thinking reportedly underperforms xhigh on 20–30% of benchmarks

keunwoochoi · x · 2026-07-25

A reply to a discussion of Opus 5 notes that Opus 4.8 does not show the same issue and suggests the behavior may indicate a rushed release.

The quoted observation underneath says that in roughly 20–30% of reported benchmarks, Opus 5 max thinking performs worse than xhigh. The surprising part is that more thinking and more test-time compute would normally be expected to improve results, so the benchmark drops may point to either an emergent issue in smaller models or a post-training problem.

Related event: Counterintuitive Benchmark: Claude Opus 5 Performs Best with Medium Reasoning(10 posts)→

Original post →

More from Models

Models channel →