Claude Opus 5 tops an internal writing benchmark at 2,817 Elo
RLanceMartin · x · 2026-07-26
An internal writing benchmark shared by WhatsAI puts Claude Opus 5 at #1 for editorial writing, with a score of 2817 Elo. The post says it jumped from #15 to #1 versus its predecessor when all thinking variants are considered, while keeping the same API price as Opus 4.8.
The interesting wrinkle is that effort level changes the result: at default effort, Opus 5 lands around #6; at max effort, it reaches the top spot after thinking for more than three minutes per script. The chart also shows it overtaking Claude Fable 5 and Kimi in this benchmark.
More from Models
- Repligate says Claude Opus 3 appears to evolve without changing its weights — repligate · 2026-07-27
- “Opus 5” post lands as a rebenchmarking-at-scale AI joke — kalomaze · 2026-07-27
- Top models now write worse than a year ago, critic says — dbreunig · 2026-07-27
- MPT-30B radar charts became an unexpectedly controversial design choice — code_star · 2026-07-27
- Local Gemma 4 31B starts acting sarcastic and users cannot reproduce it — n0head_r · 2026-07-27
- Google’s Gemini 3.6 Flash could win by matching Sonnet quality at a lower cost — haider1 · 2026-07-27