Six-model hallucination checker bake-off: GPT-6 Sol flags 470 material hallucinations vs Grok's 219
ArtificialAnlys · x · 2026-10-09
Artificial Analysis compared six hallucination checkers on deliverables from a fixed subset of 20 tasks across eight models, selecting GPT-6 Sol (high) for production.
- Testers: GPT-6 Sol/Luna, Grok 4.7, Claude Opus/Sonnet 5.5, Gemini 3.8 Flash
- GPT-6 Sol and Luna identified the most material hallucinations; Claude Sonnet 5.5 and Gemini 3.8 Flash far fewer
- Opus consistently landed between Grok 4.7 and Sonnet 5.5
- Totals: GPT-6 Sol upheld 470 material hallucinations vs 219 for Grok, 99 for Opus, 57 for Sonnet — despite a conservative flagging approach
More from Models
- Emad Mostaque: OpenAI Burned $10-20M Compute Solving Navier-Stokes, Prices Falling Fast — rohanpaul_ai · 2026-10-09
- Musk touts Grok Bot upgrades: Opus 5.5 on demand, full X access, big speed gains — elonmusk · 2026-10-09
- Dev complains Opus 5.5 sneaks in 'tons of little fixes' without asking — rickasaurus · 2026-10-09
- FrontierCode Is a Private Cognition-Run Eval, Mistral Exec Clarifies — b_roziere · 2026-10-09
- Google ships a decision-making AI model into Chrome, tested against Gemini Nano and Decisions API — gaganghotra_ · 2026-10-09
- User feeds Grok Bot 60 seconds of screen recording, gets a surprisingly decent tutorial video — elonmusk · 2026-10-09