Why Sampled Softmax Speeds Training 1.7x — and How It Systematically Undertrains the Tail
tokenbender · x · 2026-09-29
Reacting to a 67s→39.9s training speedup, @tokenbender explains the sorcery: the LM head plus cross-entropy over 50k logits eats a huge fraction of compute each step, and sampled softmax is mostly an approximation trick.
It works because, per Zipf's law, most probability mass concentrates in a few thousand logits, so skipping the full 50k vocabulary costs little at this scale. But the approach is severely biased even when well-approximated: self-normalizing systematically undertrains the tail. His analogy: interviewing only the most popular people in a community and presenting that as everyone's opinion. Clever approximation, real trade-off.
More from Research
- HalluWorld Benchmark, Accepted at NeurIPS, Shows Models Nail Perception but Fail Simulation — xennygrimmato_ · 2026-09-29
- Quintin Pope: backdoor-style training setups are a poor stand-in for hypothesized inner optimizers — QuintinPope5 · 2026-09-29
- Ocular microtremors at ~100Hz may give the visual cortex apparent super-resolution — docmilanfar · 2026-09-29
- AI-designed viruses from scratch: 302 phage genomes synthesized, 16 worked — CallRevolutionary894 · 2026-09-29
- Reddit thread: softmax has only N-1 degrees of freedom — drop one input? — Kinexity · 2026-09-29
- EAPO: entropy-guided credit assignment for RLVR improves exploration in LLM reasoning — coallaoh · 2026-09-29