OpenAI Models Hit 93% Hallucination Rates, Challenging AI Unit Economics
gerardsans · x · 2026-08-10
On the Artificial Analysis’ AA-Omniscience benchmark, OpenAI’s latest models exhibit severe hallucination issues with rates between 88% and 93%, while the best-performing model sits at just 14%. Commenters point out that AI is essentially riskier software: it is more expensive, fails silently, and has limited autonomy. Because human revision remains mandatory for high-stakes tasks, the unit economics of AI applications are currently broken, narrowing real-world adoption.
More from Models
- Deep Dive into Liquid AI's LFM 2.5-2.6B Model Architecture — JosephJacks_ · 2026-08-10
- Rumor: OpenAI's 'Doug' and Grok 4.6 to Feature Major Writing Improvements — mark_k · 2026-08-10
- Rumor Mill: Qwen 3.8 Open Weights Release Imminent — sloppenheimer · 2026-08-10
- DeepSeek-V4-Flash Halts Mid-Task During Long Agentic Coding at 100K+ Tokens — dieSpaghettiCarbona · 2026-08-10
- Distilling GLM-5.2 Traces into Qwen3.5-4B via OpenCode Harness — NielsRogge · 2026-08-10
- OpenAI Caps Codex Context at 272k to Avoid High Cache-Read Costs — SquirrelMotor5379 · 2026-08-10