Arena’s new factuality ranking puts Claude Opus 5 Max at #1
arena · x · 2026-07-28
Claude Opus 5 Max tops Arena’s new factuality ranking
Arena says Claude Opus 5 with Max reasoning is now ranked #1 on its new factuality leaderboard. The metric combines human preference with factual accuracy by sampling battles, extracting verifiable claims, and checking correctness head-to-head.
The factuality view is now live as a non-default toggle in the Text and Search Arenas, and Arena says more category-level findings will be published as more votes and traces are collected.
More from Models
- French prize-winning novel suspected of AI: $1,000 challenge over detector results — Afinetheorem · 2026-09-23
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- GPT-6 Sol priced at half of Opus 5.5 as Sol and Luna go 'dirt cheap' — ZeroStateReflex · 2026-09-23
- Tester claims Claude Opus 5.5 has the best visual design output of any model tested — burny_tech · 2026-09-23
- Meta's Alexandr Wang reveals muse has been in the works since at least Sept 2025 — adrianscottcom · 2026-09-23
- GPT-6 Sol Codex system prompt leaked: over 294,000 characters dumped on GitHub — gaganghotra_ · 2026-09-23