Renormalized scores erase confidence: KL training on them would break the model
spikedoanz · x · 2026-09-17
spikedoanz illustrates why renormalized probability scores make bad training signals: (0.09, 0.01) becomes identical to (0.9, 0.1), erasing confidence, when the model should say "idk" (0.5, 0.5). KL-training on such independently sampled scalars would likely send the model insane — a rebuttal to the claim that renormalizing merely costs a few extra ms.
Related event: Developers Warn Against Training on Renormalized Probability Scores(2 posts)→
More from Research
- CAIS ships HLE-Rolling, a continuously updated fork of Humanity's Last Exam — devindkim · 2026-09-18
- RLE-Bench shows robots can grow their own bodies — but coding agents lack physical reasoning — daibond_alpha · 2026-09-18
- ENCODE GRAMMAR models switch to interactive HTML READMEs on the ENCODE portal — anshulkundaje · 2026-09-18
- World Modeling for Physics Workshop at Aspen Center, Feb 2027 — Abstracts Due Oct 9 — randall_balestr · 2026-09-18
- Cell Focus: Biology Needs World Models That Predict Interventions, Not Just Descriptions — rishabh16_ · 2026-09-18
- RLCD: RL Fine-Tuning Pushes LLMs Toward Calibrated Confidence in Decisions — prdeepakbabu · 2026-09-18