LMSYS Launches Factuality Leaderboard to Rank AI Hallucinations

jfiance · x · 2026-08-06

LMSYS (Chatbot Arena) has announced the launch of its Factuality Leaderboards, designed to objectively evaluate the factual accuracy of AI models and guard against hallucinations.

While the traditional arena leaderboard has historically focused on human preference—and adjusted for stylistic factors like emojis and length using style control—there was a need for a more direct way to measure objective signals. The new leaderboard extracts all factual claims from a model's responses and cross-references them against the internet to ensure they are evidence-based. Models that support their claims with stronger evidence receive higher scores.

Related event: LMSYS Launches Factuality Leaderboard for AI Models(2 posts)→

Original post →

More from Models

Models channel →