Anthropic’s Opus 5 is winning benchmarks while private evals tell a different story
FinanceYF5 · x · 2026-07-28
Opus 5 is said to beat Fable on benchmarks, but private evals matter more
The post argues that Opus 5 deserves discussion because benchmark wins are no longer enough. Even though Opus 5 beats Fable on many public benchmarks, the author says people who actually use the model can tell they are not in the same league.
Three points are highlighted across the thread and images:
- Public benchmarks are losing credibility; the author now trusts private, domain-specific evals more.
- Anthropic may be trying a new training pipeline in the 5-series.
- User experience with models is getting worse overall: Claude used to feel best to work with, but that is no longer true, and the author says they now prefer Grok, with Kimi also competitive.
The thread also suggests Anthropic and OpenAI may be de-emphasizing RLHF in favor of more scalable, machine-verifiable reinforcement learning.
Related event: Anthropic Opus 5 Leads Benchmarks but Splits Real-World Reviews(6 posts)→
More from Fun
- Five Years Into the AI Boom, Google Docs Still Red-Underlines 'Compute' as a Noun — ohlennart · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- DHH Slams 'GDPR Is Good' Take: Vague Rules Birthed a Bureaucratic Beast — dhh · 2026-09-11
- Someone built a website where you can sign up for AI not to kill you — motionbynick · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Meme: Engineers Unleash 10,000 Claude Sub-Agents on Friday Afternoon to Clear a Week's Work — _jaydeepkarale · 2026-09-11