Stop Trusting Raw Softmax: Two Papers on Grounding and Calibrating LLM Output Probabilities
joecole · x · 2026-09-24
- With renewed interest in measuring LLM output probabilities sparked by jev, alex dong argues raw softmax prediction is not what you want, pointing to two papers on grounding and calibrating those probabilities.
- Quoting hxiao adds that calibration is required to make the probabilities useful for decision-making, quipping that conference-goers should stop ignoring calibration posters at ICML/ICLR/NeurIPS.
More from Research
- OpenAI Releases MentalHealthBench, an Open Benchmark Built with 80+ Clinicians — OpenAI · 2026-09-24
- MentalHealthBench Covers the Full Spectrum, From Everyday Support to Crisis Scenarios — OpenAI · 2026-09-24
- Paperena Benchmarks AI Scientists Across Full Research Cycles: Writing, Reviewing, Revising — yeewhye · 2026-09-24
- ACL'23 outstanding paper: discriminative LMs may generalize better than autoregressive models — ysu_nlp · 2026-09-24
- 10 agent reruns reached the right neighborhood, none reproduced the key observation — rohanpaul_ai · 2026-09-24
- GTSAM 4.3 ships with legged navigation, CUDA-accelerated factor graphs and continuous-time trajectory estimation — fdellaert · 2026-09-24