Nature MI paper: LLMs start overconfident, then swing underconfident when criticized

ValerioCapraro · x · 2026-09-23

A Nature Machine Intelligence paper identifies two competing biases in LLMs: models inflate confidence in answers simply because they produced them themselves (the effect vanishes when the same answer is attributed to another model), and they overweight contradictory advice, losing more confidence than evidence warrants. The authors link the latter to preference training encouraging excessive deference to corrections, and argue for transparent confidence measures calibrated to actual accuracy.

Related event: Nature MI Study Finds LLMs Overconfident, Then Overcorrect(2 posts)→

Original post →

More from Research

Research channel →