Why Models Hallucinate Instead of Admitting Ignorance: Benchmarks Reward Correctness, Not Honesty

ziv_ravid · x · 2026-07-04

The authors of the "Mirage" research explain why models tend to hallucinate rather than admit "I can't tell": it's shaped by evaluation mechanisms. Existing benchmarks only reward correct answers, not whether they are grounded. Thus, guessing based on priors has a positive expected return, while skipping yields zero—making hallucination the optimal strategy under current evaluation systems.

Original post →

More from Research

Research channel →