HalluHard Benchmark Shows Frontier Models Still Hallucinate Heavily in Multi-Turn Tasks
maksym_andr · x · 2026-09-18
New HalluHard hallucination leaderboard reveals: existing benchmarks are saturated, single-turn, and weakly judged, while HalluHard uses multi-turn tasks with rigorous full-text (PDF) verification.
- Domains: legal cases, research questions, medical guidelines, coding; toggles web search and tracks turns 1-3.
- Self-conditioning: hallucinations grow in later turns for citation tasks — 3-20% of incorrect references reappear as models condition on earlier mistakes; coding trends downward as tasks narrow.
- Capability matters: stronger models hallucinate less (GPT-5-nano → GPT-5-mini → GPT-5), with GPT-5.2 and Claude-Opus showing substantial gains.
- Reasoning isn't sufficient: effective thinking reduces hallucinations for GPT-family models but is model-dependent (no improvement for DeepSeek-Reasoner).
Frontier proprietary models struggle even with web search enabled.
Related event: GPT-6-Astra Tops HalluHard Benchmark as Multi-Turn Hallucinations Persist(3 posts)→
More from Models
- Stanford's 10-person Marin open lab is live-training a 535B model in the open — wandb · 2026-09-18
- Self-Proclaimed ChatGPT Co-Inventor Launches Jev, Claims 200x Speed at 1/400 Cost — iamrobotbear · 2026-09-18
- Codex Pro User Says Usage Limits Got 5-10x Worse, Can't Even Buy Another Plan — Junra · 2026-09-18
- GPT-6 Astra Deciphers an Undeciphered 1918 German WWI Radio Transmission — moultano · 2026-09-18
- Sakana AI Introduces Fugu Max and Fugu Ultra v2 Models — SakanaAILabs · 2026-09-18
- Gemini 3.8 Live Architecture Breakdown: Sub-100ms Native Audio and Real-Time Tool Calling — 4bTechDecode · 2026-09-18