HalluHard Benchmark: GPT-6-Astra Clearly Ahead of Fable 5 on Hallucinations

maksym_andr · x · 2026-09-18

New results on the HalluHard hallucination benchmark: GPT-6-Astra is significantly better than any other model, including Fable 5, both with and without web search. The author notes few people discuss hallucinations now, but they remain an important indicator of lack of uncertainty awareness.

Related event: GPT-6-Astra Tops HalluHard Benchmark as Multi-Turn Hallucinations Persist(3 posts)→

Original post →

More from Models

Models channel →