Claude Opus 5 may be scoring lower on hard benchmarks because it offloads work to subagents

draginol · x · 2026-07-25

A comment suggests Claude Opus 5 may be underperforming on “higher effort” benchmarks because it delegates more work to subagents, and those subagents are often not Opus 5 itself.

Related event: Counterintuitive Benchmark: Claude Opus 5 Performs Best with Medium Reasoning(10 posts)→

Original post →

More from Models

Models channel →