One confidence signal controls LLM abstention: 66.5% to 7% in Gemma 3 27B
dejanseo · x · 2026-09-08
A study by Kumaran et al. in Nature Machine Intelligence finds that LLM abstention is governed by a single internal confidence signal that outweighs question difficulty, fact retrievability, and wording by roughly 10x.
- Tested on GPT-4o, Gemma 3 27B, DeepSeek 671B and Qwen 80B; Llama 3.1 70B was dropped for low abstention
- GPT-4o baseline: 63.7% accuracy with forced response; calibrated confidence tracked errors at r = -0.97
- With an abstention option, GPT-4o abstained on 56.6% of items, raising answered accuracy to 69.1%; abstention ranged 27-82% across models
- Activation steering in Gemma 3 27B swung abstention from 66.5% to 7.0% (28.2% baseline), with the signal mediating 67.1% of the effect
- Injecting a high-confidence signal could make models answer questions they would otherwise skip
More from Models
- Blogger: Of course OpenAI reads your 'private' Codex logs and stealth-nerfs models — Linahuaa · 2026-09-08
- GLM 5.3 Flash vs DeepSeek V4.1: voxel scene test ends in a draw despite 1.5x more steps — teortaxesTex · 2026-09-08
- Blogger's internal benchmarks: Chinese AI models pricier than US models quality-adjusted — kevinnbass · 2026-09-08
- Buy an M5 Max Mac Studio (128GB, €5,849) for local Qwen now, or wait for M7 Ultra in 2028? — Mxmtm · 2026-09-08
- DeepSeek Flash 4.1 spotted testing via API with new architecture, release imminent — kimmonismus · 2026-09-08
- Zvi: OpenAI's Astra Marks a Rapid Decline in Chain-of-Thought Monitorability — Don't Worry About the Vase (Zvi) · 2026-09-08