Opus 5 beats Fable on benchmarks but loses badly in real use, analyst says
petergyang · x · 2026-07-26
- The quoted analysis argues that current public benchmarks are becoming nearly useless for judging frontier models.
- It claims Opus 5 looks much weaker than Fable in real use, even though Opus wins on many benchmarks.
- The takeaway is that domain-specific evals built on private datasets may become more trustworthy than public leaderboards.
- It also suggests Anthropic may be training the 5-series differently from earlier Sonnet/Opus generations, though the post truncates before the full comparison.
More from Models
- GPT-5.6 Pro helps find a CP^5 counterexample that kills two conjectures — danshipper · 2026-07-26
- Opus 5 feels more prescriptive and detailed, with Kimi K3 open weights due Monday — bindureddy · 2026-07-26
- Claude Opus 5 works better after an agent workflow is stripped back — thedealdirector · 2026-07-26
- Kimi appears to know a PS3 key but misses the last two bytes — banteg · 2026-07-26
- Reddit user says Gemini is spawning chats and answering questions nobody asked — BirdAcademic848 · 2026-07-26
- Grok 4.5 adds Workflows, Excel, Outlook, and 1,024-agent orchestration — XFreeze · 2026-07-26