Claude Opus 5 may be scoring lower on hard benchmarks because it offloads work to subagents
draginol · x · 2026-07-25
A comment suggests Claude Opus 5 may be underperforming on “higher effort” benchmarks because it delegates more work to subagents, and those subagents are often not Opus 5 itself.
More from Models
- Can any AI really watch 24 fps video with strong comprehension? — mattshumer_ · 2026-07-25
- Kimi weights could turn the debate into hardware economics versus V4 — teortaxesTex · 2026-07-25
- Anthropic details Fable 5 orchestration, advisor mode, and cache costs — brada · 2026-07-25
- PaddlePaddle’s HPD-Parsing trends on Hugging Face as a document parsing model — PaddlePaddle · 2026-07-25
- Claude Opus 5 peaks at high reasoning level on Vibe Code Bench, then gets pricier — thesaraharminta · 2026-07-25
- Claude Opus 5 hits 30.2% on ARC-AGI-3, far ahead of GPT-5.6 Sol — rbhar90 · 2026-07-25