Free models aren't the real problem: studies show hallucinations persist in SOTA LLMs
AryHHAry · x · 2026-09-19
Responding to a claim that AI skeptics are just using dumb free tiers, the author argues paid/reasoning models don't fix hallucinations — it's a statistical property of training and evaluation pipelines that reward confident guesses over admissions of ignorance.
Cited evidence:
- Kalai et al. (OpenAI + Georgia Tech, 2025): hallucinations persist even in SOTA systems.
- Dahl et al. (Journal of Legal Analysis, 2024): public models including GPT-4 hallucinate on verifiable legal questions.
- Magesh et al. (Stanford, 2024): paid legal research tools like Lexis and Westlaw still err 17–33% of the time.
The author notes high-stakes domains like legal documents and code are precisely where hallucinations need to be measured, rather than dismissing skeptics as users of cheap models.
More from Models
- Rumor: Anthropic's Opus 5.2 and Fable 5.2 already tested, release may slip with IPO — kimmonismus · 2026-09-19
- Probability-Only Model Jev Beats or Matches GPT-5.6-luna on 42 of 49 Tasks at 4x Speed, 1/4 Cost — LowNefariousness9966 · 2026-09-19
- Codex usage limits finally reset after days of users hitting exhausted quotas — CtrlAltDwayne · 2026-09-19
- Rumored Opus 5.2 generates a full 5-minute Titanic film in pure JavaScript with three.js — dotey · 2026-09-19
- Model Remembers 'cranberry-42' Across App Restarts in Persistent-Session Test — RileyRalmuto · 2026-09-19
- Why do AI models lack personality? Safety training and attachment fears, explained — alexisgallagher · 2026-09-19