Study: 5 Frontier Models Disagree on 23% of Fact-Check Claims

Pavel___1__ · reddit · 2026-08-25

A study by Lenz tested 1,000 real-world fact-check claims across five frontier models (Claude Fable 5, GPT-5.6, Gemini 3.1 Pro, Sonar Deep Research, Grok 4.5) with web access enabled. Results show 37% total agreement, but 23% 'material disagreement' (2+ steps apart on a 5-point scale). Distinct 'personalities' emerged: Gemini is polarized (True/False), Sonar constantly hedges, and Claude matches the majority verdict 86% of the time. Crucially, self-reported confidence is near-useless: 76% of all answers were rated 9/10 or 10/10, even on claims where models collectively disagreed. The author warns that a single model's verdict is vendor-dependent, and confidence scores are poor signals for human escalation compared to inter-model disagreement. Code and data are open-sourced.

Original post →

More from Models

Models channel →