Study: 5 Frontier Models Disagree on 23% of Fact-Check Claims
Pavel___1__ · reddit · 2026-08-25
A study by Lenz tested 1,000 real-world fact-check claims across five frontier models (Claude Fable 5, GPT-5.6, Gemini 3.1 Pro, Sonar Deep Research, Grok 4.5) with web access enabled. Results show 37% total agreement, but 23% 'material disagreement' (2+ steps apart on a 5-point scale). Distinct 'personalities' emerged: Gemini is polarized (True/False), Sonar constantly hedges, and Claude matches the majority verdict 86% of the time. Crucially, self-reported confidence is near-useless: 76% of all answers were rated 9/10 or 10/10, even on claims where models collectively disagreed. The author warns that a single model's verdict is vendor-dependent, and confidence scores are poor signals for human escalation compared to inter-model disagreement. Code and data are open-sourced.
More from Models
- Test of 1,000 prompts: ChatGPT and Google AI share only 2.6% of cited sources — nikvassev · 2026-08-25
- Rumor: Alibaba's Qwen to release a 125B-A6B MoE model, hailing the era of ~120B models — TheZachMueller · 2026-08-25
- Unsloth AI aims for day-zero llama.cpp support for Qwen models — danielhanchen · 2026-08-25
- MiniMax Launches H3 Model and End-to-End Design Platform — mhdfaran · 2026-08-25
- Qwen Announces Qwen3.8-Flash-Next Open Source Release — cedric_chee · 2026-08-25
- AI detection tool flags content as likely generated by Claude — pedrodias · 2026-08-25