Arena Introduces Factuality Leaderboard
Kyrannio · x · 2026-07-17
Arena has introduced a new model ranking dimension: factuality.
- Rankings are no longer based solely on human preference; instead, human preference and factuality are combined into a weighted metric. This new view is available as a non-default toggle in Text Arena and Search Arena.
- Their approach involves randomly sampling battles, extracting web-verifiable claims from model responses, verifying each claim for accuracy, and comparing the average accuracy across different models.
- To support this leaderboard, Arena has annotated over 2 million LLM claims, including 1.3 million+ from Text Arena and 700,000+ from Search Arena.
- A related video explains how this method extends the Bradley-Terry objective to account for both preference and factuality, and why models shouldn't be penalized simply for "refusing to answer / making no assertions."
Related event: Arena Adds Factuality to Model Rankings(12 posts)→
More from Research
- Cognition's SWE-2 uses a KKT duality argument in RL to shift the effort Pareto curve — YouJiacheng · 2026-09-11
- VidMap uses RoMa coarse matching on all frames, fine-scale only for keyframes — ducha_aiki · 2026-09-11
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11