Arena Adds Factuality to Model Rankings

Arena announced on July 16 that it is adding factuality to its model rankings in both Text Arena and Search Arena. The new view combines human preference and factuality through adjustable weights and is offered as a non-default toggle. The update matters because it tries to separate “what users like” from “what is actually correct,” then compare models on both axes together.

Method and data

According to Arena, the factuality signal is built by randomly sampling battle outputs, extracting claims that can be verified on the web, and then checking those claims one by one. To support this system, Arena said it has annotated more than 2 million claims from real LLM conversations, including more than 1.3 million from Text Arena and more than 700,000 from Search Arena. The team also published statistics on how many claims appear and how they are distributed across battles.

Ranking changes

In Search Arena, turning on factuality changes the standings: GPT-5.5-search rises to the top, while GPT-5.2-search jumps from No. 11 to No. 3. Arena also released an interactive weighted chart so users can inspect how scores and rankings move as the factuality weight changes. In a separate observation, Arena said most open-source models lose score as factuality weight increases, with nvidia-nemotron-3-ultra dropping especially sharply; mistral-medium-3.5 was cited as one of the exceptions.

Why it matters

Arena said the metric reveals differences among frontier models that were harder to see in preference-only rankings and better matches real-world expectations around reliability. In practice, this is not just a new filter for the leaderboard, but an attempt to evaluate usefulness and correctness separately before recombining them.

2026-07-16 ~ 2026-07-17 · 12 related posts

6 near-duplicate retellings: arena · arena · the-grand-finale · WeijiaShi2 · CSProfKGD · Kyrannio