Why Sampled Softmax Speeds Training 1.7x — and How It Systematically Undertrains the Tail

tokenbender · x · 2026-09-29

Reacting to a 67s→39.9s training speedup, @tokenbender explains the sorcery: the LM head plus cross-entropy over 50k logits eats a huge fraction of compute each step, and sampled softmax is mostly an approximation trick.

It works because, per Zipf's law, most probability mass concentrates in a few thousand logits, so skipping the full 50k vocabulary costs little at this scale. But the approach is severely biased even when well-approximated: self-normalizing systematically undertrains the tail. His analogy: interviewing only the most popular people in a community and presenting that as everyone's opinion. Clever approximation, real trade-off.

Original post →

More from Research

Research channel →