Arena’s new factuality ranking puts Claude Opus 5 Max at #1
arena · x · 2026-07-28
Claude Opus 5 Max tops Arena’s new factuality ranking
Arena says Claude Opus 5 with Max reasoning is now ranked #1 on its new factuality leaderboard. The metric combines human preference with factual accuracy by sampling battles, extracting verifiable claims, and checking correctness head-to-head.
The factuality view is now live as a non-default toggle in the Text and Search Arenas, and Arena says more category-level findings will be published as more votes and traces are collected.
More from Models
- Claude Opus 5 Max takes No. 1 in Arena’s frontend coding leaderboard — scaling01 · 2026-07-28
- Continual learning on Qwen3.5-397B is said to match Opus 4.8 for about $450k — josh_wills · 2026-07-28
- ProteinGym-LLM ranks Claude Opus 5 highest on a 217-task protein variant benchmark — LeoTZ03 · 2026-07-28
- Kimi K2 Third-Party Inference Pricing Matches Official API as Location Becomes Key — kevinsxu · 2026-07-28
- Kimi K3 used 51.2 million sandboxes across 1.5 million images — tarantulae · 2026-07-28
- Developer Rants: Current SOTA Models Are Practically Worse Than Last Gen — zeeg · 2026-07-28