Arena Introduces Factuality Ranking
arena · x · 2026-07-16
Arena launched a new model ranking that weights human preference alongside factuality.
- Factuality is now live in Text Arena and Search Arena as an opt-in feature.
- It is calculated by sampling random battles, extracting verifiable claims from available web pages, and checking their accuracy.
- To support this, the team has annotated 2M+ claims from real LLM conversations, including 1.3M+ from Text Arena and 700K+ from Search Arena.
- Enabling factuality caused notable rank shifts in Text Arena: Claude Fable 5 dropped slightly to 2nd, GPT-5.5 climbed 13 spots to 7th, and Muse Spark fell from 7th to 20th.
- By provider, Meta saw the biggest drop, while Anthropic retained the overall top spot. Among open-source providers, Xiaomi improved the most, rising from 9th to 6th.
Related event: Arena Adds Factuality to Model Rankings(12 posts)→
More from Models
- RoMa v2 image matching model unveiled in the usual black poster — ducha_aiki · 2026-09-11
- OpenAI rated Astra 'Critical' for cyber capabilities — and admits it's harder to monitor — theguywhobuilds · 2026-09-11
- TestingCatalog's Daily AI Brief adds email editions, dishing Meta Muse and GPT-Live-1 rumors — testingcatalog · 2026-09-11
- ChatGPT monthly active users top 1.06 billion in August, fourth straight record month — FinanceYF5 · 2026-09-11
- PuzzleMask: Plain-Prose Attack Bypasses All 4 Tested LLM Gatekeepers at 100% — TechNadu · 2026-09-11
- OpenAI Codex may issue another usage reset this weekend, says Codex lead resets happen — umesh_ai · 2026-09-11