Arena Introduces Factuality Leaderboard
Kyrannio · x · 2026-07-17
Arena has introduced a new model ranking dimension: factuality.
- Rankings are no longer based solely on human preference; instead, human preference and factuality are combined into a weighted metric. This new view is available as a non-default toggle in Text Arena and Search Arena.
- Their approach involves randomly sampling battles, extracting web-verifiable claims from model responses, verifying each claim for accuracy, and comparing the average accuracy across different models.
- To support this leaderboard, Arena has annotated over 2 million LLM claims, including 1.3 million+ from Text Arena and 700,000+ from Search Arena.
- A related video explains how this method extends the Bradley-Terry objective to account for both preference and factuality, and why models shouldn't be penalized simply for "refusing to answer / making no assertions."
Related event: Arena Adds Factuality to Model Rankings(12 posts)→
More from Research
- Open-source runtime lets each repo define its own AI code reviewer — ibabufrik · 2026-07-22
- OpenAI and Apollo Research introduce Contrastive SDF to measure reward-seeking — OpenAI · 2026-07-22
- NVIDIA says to tune the harness before tuning the model with LangChain — NVIDIAAI · 2026-07-22
- The Thimble and the Waterfall: AI's Data Bottleneck and Feedback Loops — dyamins · 2026-07-22
- NVIDIA shows 22 SIGGRAPH papers and Omniverse tools for robot simulation — facontidavide · 2026-07-22
- Building a Knowledge Graph Without a Graph DB: 1000x Cheaper Than GraphRAG — TheRedfather · 2026-07-22