Renormalized scores erase confidence: KL training on them would break the model

spikedoanz · x · 2026-09-17

spikedoanz illustrates why renormalized probability scores make bad training signals: (0.09, 0.01) becomes identical to (0.9, 0.1), erasing confidence, when the model should say "idk" (0.5, 0.5). KL-training on such independently sampled scalars would likely send the model insane — a rebuttal to the claim that renormalizing merely costs a few extra ms.

Related event: Developers Warn Against Training on Renormalized Probability Scores(2 posts)→

Original post →

More from Research

Research channel →