Reasoning Models Stay Miscalibrated: EMNLP Studies on AI Confidence
zhaoran_wang · x · 2026-09-29
Sharing work sparked by Jev/RLCD putting calibration in the spotlight, the authors report: two EMNLP 2025 papers evaluating calibration in reasoning models and VLMs found stronger capability doesn't automatically mean more reliable confidence — task and modality matter. Their ACL 2026 work extends this to agentic RL, explicitly optimizing calibration alongside task performance in tool-using agents ("The Confidence Dichotomy"), arguing calibration becomes a core capability as confidence starts driving real actions.
More from Research
- 30+ labs fail to replicate Marcus et al's 1999 Science paper on infants vs RNNs, built on just 16 babies — tallinzen · 2026-09-29
- Chris Manning on why LMs learn verb categories first, and why he thinks LeCun is wrong about language — ziv_ravid · 2026-09-29
- Researchers clash over what a valid Bayesian updating rule in research synthesis should look like — RexDouglass · 2026-09-29
- Google Research unveils multi-agent 'AI video co-director' for consistent long-form video generation — rseroter · 2026-09-29
- BLUE from Google internship lands NeurIPS: LLM-written user profiles boost recommendations — shangbinfeng · 2026-09-29
- Collatz twist: Krasikov-Lagarias-style X^0.84 bounds apply to any root, and equally to 3x-1 — AlexKontorovich · 2026-09-29