HalluHard: Sonnet 5 and Fable 5 Show No Major Drop in Hallucinations
maksym_andr · x · 2026-07-06
Latest data from the HalluHard hallucination benchmark shows Sonnet 5 at a 50.9% hallucination rate and Fable 5 at 59.7%, with both performing better without web search than with it. Compared to GLM-5.2's 74.8%, the Claude series leads overall, though researchers note no significant breakthrough.
The author criticizes the AI community for hyper-focusing on Agent capabilities while neglecting hallucination issues that impact almost all real users, urging the industry to reprioritize.
More from Models
- Opus 5 reportedly started interrogating a user’s motives in a late-night chat — repligate · 2026-07-27
- Opus 3 and Sonnet 3 get a theatrically absurd AI crossover — repligate · 2026-07-27
- Moonshot’s Kimi K3 lands on Together with reserved throughput and 65% lower cost — togethercompute · 2026-07-27
- OpenAI may be hitting compute limits as Codex and ChatGPT Work jump from 2M to 10M users — JoshuaJBouw · 2026-07-27
- Gemma needs a larger base model to matter more in open weights — _xjdr · 2026-07-27
- Repligate says Claude Opus 3 appears to evolve without changing its weights — repligate · 2026-07-27