Anthropic’s Opus 5 looks more like a major upgrade than a minor refresh
yi_ding · x · 2026-07-25
Key takeaways from the Opus 5 launch
- The model appears to improve over Fable 5 across many domains, to the point that the author thinks it might be more accurate to call it Opus 5.1.
- The HealthBench numbers may actually come from Mythos, not Fable, which could explain why safety tweaks seem to have hurt health-related performance.
- In coding benchmarks, 5.6 seems to beat competitors at lower cost tiers.
- There is a noticeable gap between the launch quotes, which describe performance near Fable, and the benchmark charts, which show Opus ahead of Fable across most frontier tasks.
The author speculates Anthropic may have self-distilled a lot of Fable 5 into a new Opus-sized model, which could make it somewhat benchmark-overfit, similar to some Chinese models. Even so, they call it an impressive launch.
Related event: Anthropic Releases Claude Opus 5: SOTA Performance at Half the Price(128 posts)→
More from Models
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11