Anthropic’s Opus 5 is winning benchmarks while private evals tell a different story
FinanceYF5 · x · 2026-07-28
Opus 5 is said to beat Fable on benchmarks, but private evals matter more
The post argues that Opus 5 deserves discussion because benchmark wins are no longer enough. Even though Opus 5 beats Fable on many public benchmarks, the author says people who actually use the model can tell they are not in the same league.
Three points are highlighted across the thread and images:
- Public benchmarks are losing credibility; the author now trusts private, domain-specific evals more.
- Anthropic may be trying a new training pipeline in the 5-series.
- User experience with models is getting worse overall: Claude used to feel best to work with, but that is no longer true, and the author says they now prefer Grok, with Kimi also competitive.
The thread also suggests Anthropic and OpenAI may be de-emphasizing RLHF in favor of more scalable, machine-verifiable reinforcement learning.
Related event: Anthropic's Opus 5 Wins Benchmarks but Faces Private Testing Backlash(2 posts)→
More from Fun
- National Crème Brûlée Day AI challenge showcase thread goes live — LudovicCreator · 2026-07-28
- A game character walking a dog turns into a cute multimodal mashup — Humble-Tangerine-878 · 2026-07-28
- A user says Opus 5 made 17 monthly exports unnecessary in a coding workflow — gaganghotra_ · 2026-07-28
- An AI-made live-action One Piece is the version fans say they actually wanted — eyishazyer · 2026-07-28
- Musk says compute can be revoked from AI firms that “harm humanity” — XFreeze · 2026-07-28
- A puppy stays unusually calm around Google OmniFlash — jnack · 2026-07-28