Nature MI paper: LLMs start overconfident, then swing underconfident when criticized
ValerioCapraro · x · 2026-09-23
A Nature Machine Intelligence paper identifies two competing biases in LLMs: models inflate confidence in answers simply because they produced them themselves (the effect vanishes when the same answer is attributed to another model), and they overweight contradictory advice, losing more confidence than evidence warrants. The authors link the latter to preference training encouraging excessive deference to corrections, and argue for transparent confidence measures calibrated to actual accuracy.
Related event: Nature MI Study Finds LLMs Overconfident, Then Overcorrect(2 posts)→
More from Research
- ICLR load debate: capping submissions per author would barely dent paper counts, data shows — jonasgeiping · 2026-09-23
- Trust AI R&D evals only if scored by people who've hand-labeled outputs — dfrsrchtwts · 2026-09-23
- JevBench Model Decision Index Debuts on Hugging Face, Native Image Support Still Missing — openSourcerer9000 · 2026-09-23
- AI swarms spontaneously grow scale-free hub topologies, MIT professor observes — ProfBuehlerMIT · 2026-09-23
- Flex-π: a 6B world-action model beats π0.5 by up to 6x on real bimanual tasks — chris_j_paxton · 2026-09-23
- Deoptimization Is Harder Than Optimization: How JIT Escape Hatches Keep Dynamic Languages Fast — lauriewired · 2026-09-23