Raw LLM probabilities aren't enough for decisions — calibration matters, researchers argue
PMinervini · x · 2026-09-23
@tinkerapi argued that LLM next-token prediction is already a probabilistic classifier, so an open LLM can serve a Jev-like interface with fast probability outputs, improvable with a $5, 10-minute Tinker run. @hxiao pushed back: raw probabilities need calibration to be useful for decision-making, urging people not to skip calibration posters at ICML/ICLR/NeurIPS.
More from Models
- Beff Jezos jokes Opus 5.5 is 'post-slop', freeing readers from AI sludge prose — beffjezos · 2026-09-23
- Opus 5.5 One-Shots a Full Prince of Persia Level with NPCs, Sound and Music — iannuttall · 2026-09-23
- The Full Prompt Behind Opus 5.5's One-Shot Prince of Persia Level — iannuttall · 2026-09-23
- Model release fatigue: devs say gains are marginal, but some argue new models clearly outpace old ones — Rasmic · 2026-09-23
- Follow-up: Opus 5.5 roughly on par with Astra and Fable 5.1, no clear winner — AaronBergman18 · 2026-09-23
- Opus 5.5 tentatively the world's smartest model, though slightly behind on world knowledge — AaronBergman18 · 2026-09-23