DeepMind-Princeton paper shows LLMs causally use confidence to decide whether to answer
GoogleDeepMind · x · 2026-09-07
Researchers from Google DeepMind and Princeton (lead author Dharshan Kumaran) published an open-access Nature Machine Intelligence paper offering causal evidence that LLMs use confidence signals to drive behavior, not just passively report them.
Key points:
- A four-phase paradigm elicited baseline confidence without abstention, then probed abstention behavior
- LLMs apply an implicit threshold on internal confidence when abstaining; confidence effect sizes were roughly an order of magnitude larger than alternative mechanisms
- Causal evidence via activation steering: boosting confidence decreased abstention, suppressing it increased abstention, confirmed by mediation analysis
Conclusion: models genuinely use their own confidence to decide whether to answer or abstain, paralleling metacognition in biological systems.
More from Models
- Mystery model Omen Alpha spotted; tokenizer tests point to new Zhipu GLM — realsohamparekh · 2026-09-07
- Training mixtures are now all synthetic: small-model training is really distillation — RexDouglass · 2026-09-07
- Sol High usage test: one complex prompt eats 5% of the 5-hour limit — remixedmoon5 · 2026-09-07
- Bodhan AI open-weights speech, vision and translation models for Indian languages on Hugging Face — selfawareatom · 2026-09-07
- Claude Max user reports a week of erratic usage-limit bugs and resets — tonimedic · 2026-09-07
- Dev Reminder: Astra Shines in Demo-Friendly Domains, but AGI Hinges on System-Level Understanding — Scobleizer · 2026-09-07