HalluHard crowns new leader: GPT-6-Astra beats all models on hallucination
maksym_andr · x · 2026-09-18
HalluHard's leaderboard has a new #1: GPT-6-Astra significantly outperforms every other model, including Fable 5, with and without web search. The author argues hallucination remains a key indicator of poor uncertainty awareness even as overall rates trend down.
Related event: GPT-6-Astra Tops HalluHard Hallucination Benchmark(4 posts)→
More from Models
- Google's Astra plays Unciv faster than humans by batching moves — Angaisb_ · 2026-09-18
- Mozilla report: open-weight models 4 months behind frontier, Qwen hits 942M downloads — rohanpaul_ai · 2026-09-18
- DeepSeek releases V4.1-Flash: smallest model in new family with native vision, $0.15/M tokens — aziz4ai · 2026-09-18
- What GPT-6 Astra's 99.9% ARC-AGI-3 Score Actually Measures — mixtapedmonk · 2026-09-18
- Unreleased Astra-family model got a new persona in RL training, and people love it — cephaloform · 2026-09-18
- Embedded messages shape training far more than inference — they end up in the weights — Gccooke · 2026-09-18