Search Arena Factuality Leaderboard Updated
arena · x · 2026-07-16
Search Arena has announced the latest leaderboard changes related to factuality, providing interactive weighted charts and methodology notes to visualize how scores and rankings shift under different weights.
Key changes mentioned include:
- GPT-5.5-search moved up 1 spot to take the #1 position in factuality
- GPT-5.2-search saw a massive jump from #11 to #3
- claude-sonnet-4-6-search dropped from #6 to #9
- gemini-3.1-pro-grounding fell from #7 to #13
Related event: Arena Adds Factuality to Model Rankings(12 posts)→
More from Models
- Daily AI brief: GPT-Live-1 in API, OpenAI pauses $200 Pro signups amid Astra demand — koltregaskes · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11