Arena Leaderboard Adds Factuality Metrics

CSProfKGD · x · 2026-07-16

Arena announced that its Text and Search leaderboards now employ a hybrid ranking of human preference + factuality, with factuality featured as a toggleable, non-default option.

They evaluate the accuracy of model responses by randomly sampling battle instances, extracting verifiable web-based claims, and verifying them one by one. To support this ranking system, Arena has annotated over 2 million model claims from real-world conversations, including over 1.3 million from Text Arena and over 700,000 from Search Arena.

The post also notes that when factuality is enabled, some models see significant ranking shifts: OpenAI's models are largely unaffected; Grok improves due to its emphasis on truth and objectivity; Anthropic and Google show mixed results; Meta's Muse Spark and Dola Seed drop significantly, though Muse Spark 1.1 still outperforms version 1.0 and the Llama series.

Related event: Arena Adds Factuality to Model Rankings(12 posts)→

Original post →

More from Models

Models channel →