HalluHard Benchmark Shows Frontier Models Still Hallucinate Heavily in Multi-Turn Tasks

maksym_andr · x · 2026-09-18

New HalluHard hallucination leaderboard reveals: existing benchmarks are saturated, single-turn, and weakly judged, while HalluHard uses multi-turn tasks with rigorous full-text (PDF) verification.

Frontier proprietary models struggle even with web search enabled.

Related event: GPT-6-Astra Tops HalluHard Benchmark as Multi-Turn Hallucinations Persist(3 posts)→

Original post →

More from Models

Models channel →