HalluHard Results: GPT-6-Astra Beats Fable 5 and All Others on Hallucination Control
maksym_andr · x · 2026-09-18
New results on the HalluHard hallucination benchmark show GPT-6-Astra significantly outperforming every other model, including Fable 5, both with and without web search. The author argues hallucinations are under-discussed lately but remain a key indicator of models' lack of uncertainty awareness. This follow-up adds the no-web-search results.
Related event: GPT-6-Astra Tops HalluHard Benchmark as Multi-Turn Hallucinations Persist(3 posts)→
More from Models
- Stanford's 10-person Marin open lab is live-training a 535B model in the open — wandb · 2026-09-18
- Self-Proclaimed ChatGPT Co-Inventor Launches Jev, Claims 200x Speed at 1/400 Cost — iamrobotbear · 2026-09-18
- Codex Pro User Says Usage Limits Got 5-10x Worse, Can't Even Buy Another Plan — Junra · 2026-09-18
- GPT-6 Astra Deciphers an Undeciphered 1918 German WWI Radio Transmission — moultano · 2026-09-18
- Sakana AI Introduces Fugu Max and Fugu Ultra v2 Models — SakanaAILabs · 2026-09-18
- Gemini 3.8 Live Architecture Breakdown: Sub-100ms Native Audio and Real-Time Tool Calling — 4bTechDecode · 2026-09-18