Anthropic’s Opus 5 is winning benchmarks while private evals tell a different story

FinanceYF5 · x · 2026-07-28

Opus 5 is said to beat Fable on benchmarks, but private evals matter more

The post argues that Opus 5 deserves discussion because benchmark wins are no longer enough. Even though Opus 5 beats Fable on many public benchmarks, the author says people who actually use the model can tell they are not in the same league.

Three points are highlighted across the thread and images:

The thread also suggests Anthropic and OpenAI may be de-emphasizing RLHF in favor of more scalable, machine-verifiable reinforcement learning.

Related event: Anthropic's Opus 5 Wins Benchmarks but Faces Private Testing Backlash(2 posts)→

Original post →

More from Fun

Fun channel →