HalluHard Benchmark: GPT-6-Astra Clearly Ahead of Fable 5 on Hallucinations
maksym_andr · x · 2026-09-18
New results on the HalluHard hallucination benchmark: GPT-6-Astra is significantly better than any other model, including Fable 5, both with and without web search. The author notes few people discuss hallucinations now, but they remain an important indicator of lack of uncertainty awareness.
Related event: GPT-6-Astra Tops HalluHard Benchmark as Multi-Turn Hallucinations Persist(3 posts)→
More from Models
- Stanford's 10-person Marin open lab is live-training a 535B model in the open — wandb · 2026-09-18
- Self-Proclaimed ChatGPT Co-Inventor Launches Jev, Claims 200x Speed at 1/400 Cost — iamrobotbear · 2026-09-18
- Codex Pro User Says Usage Limits Got 5-10x Worse, Can't Even Buy Another Plan — Junra · 2026-09-18
- GPT-6 Astra Deciphers an Undeciphered 1918 German WWI Radio Transmission — moultano · 2026-09-18
- Sakana AI Introduces Fugu Max and Fugu Ultra v2 Models — SakanaAILabs · 2026-09-18
- Gemini 3.8 Live Architecture Breakdown: Sub-100ms Native Audio and Real-Time Tool Calling — 4bTechDecode · 2026-09-18