Raw LLM probabilities aren't enough for decisions — calibration matters, researchers argue

PMinervini · x · 2026-09-23

@tinkerapi argued that LLM next-token prediction is already a probabilistic classifier, so an open LLM can serve a Jev-like interface with fast probability outputs, improvable with a $5, 10-minute Tinker run. @hxiao pushed back: raw probabilities need calibration to be useful for decision-making, urging people not to skip calibration posters at ICML/ICLR/NeurIPS.

Original post →

More from Models

Models channel →