Search Scenario Factuality Rankings Updated
arena · x · 2026-07-16
Arena provided an update on the factuality leaderboard changes in Search Arena.
- With factuality enabled, GPT-5.5-search climbed to the top spot.
- GPT-5.2-search saw a more significant leap, moving from #11 to #3.
- claude-sonnet-4-6-search dropped from #6 to #9, and gemini-3.1-pro-grounding fell from #7 to #13.
- They also noted that while some open-source models saw score declines when factuality weights were increased, mistral-medium-3.5 and tencent huyuan-hy3-preview actually moved up.
Related event: Arena Adds Factuality to Model Rankings(12 posts)→
More from Models
- OpenAI is reportedly briefing U.S. lawmakers on its next model family — kimmonismus · 2026-07-22
- Muse Spark 1.1 lands at 1495 on Text Arena with standout agentic-coding price performance — ycombinator · 2026-07-22
- Advanced AI Models Are Becoming Impossible to Plug and Play — emollick · 2026-07-22
- Google Gemini's AI Problem: No Leading Model for Core Workloads — bindureddy · 2026-07-22
- Model Offers 1M Token Context Window at Just $0.33/1M Tokens — MickeySteamboat · 2026-07-22
- Google says information agents are coming to AI Pro and Ultra this summer — gaganghotra_ · 2026-07-22