Anthropic’s Opus 5 is winning benchmarks while private evals tell a different story

FinanceYF5 · x · 2026-07-28

Opus 5 is said to beat Fable on benchmarks, but private evals matter more

The post argues that Opus 5 deserves discussion because benchmark wins are no longer enough. Even though Opus 5 beats Fable on many public benchmarks, the author says people who actually use the model can tell they are not in the same league.

Three points are highlighted across the thread and images:

The thread also suggests Anthropic and OpenAI may be de-emphasizing RLHF in favor of more scalable, machine-verifiable reinforcement learning.

Related event: Anthropic Opus 5 Leads Benchmarks but Splits Real-World Reviews(6 posts)→

Original post →

More from Fun

Fun channel →