Softmax picks probabilities, cross-entropy picks the target: a 3-class walkthrough
techNmak · x · 2026-10-11
A clear walkthrough of the classifier tail using logits 2, 1, 0 where the correct class C has the lowest score:
- Different jobs: Softmax only converts logits into a probability distribution (66.5% / 24.5% / 9.0%); cross-entropy brings in the target. Loss depends solely on the probability assigned to the correct class — here just 9%, giving loss ≈ 2.4076.
- Log penalty: High probability on the correct class means small loss; as it approaches zero, loss grows without bound.
- Simple gradient: The gradient w.r.t. logits is just predicted probability minus target: +0.665 and +0.245 for wrong classes, -0.910 for the correct one, then backprop carries these through the network.
- Numerical stability: Subtract the max logit before exponentiating to avoid overflow; probabilities and loss are unchanged.
- Practical tip: PyTorch's crossentropy expects raw logits — don't apply softmax first.
More from Research
- Inferring goals from failure: online Bayesian goal inference for boundedly-rational agents — xuanalogue · 2026-10-11
- Alignment researcher points to Rohin Shah's value learning sequence and IRL model misspecification — xuanalogue · 2026-10-11
- Looped LM paper: 1.6B model matches full-cache baseline with 3x smaller KV cache — rupspace · 2026-10-11
- Pure RL discovers superhuman robot strategies in sim, transfers zero-shot to real hardware — KyleMorgenstein · 2026-10-11
- PartLLM brings LLM-powered 3D mesh part segmentation to SIGGRAPH Asia with code released — Promptmethus · 2026-10-11
- Duo Bregman pseudo-divergence yields closed-form KL divergence between truncated Gaussians — FrnkNlsn · 2026-10-11