Paper: LM Head Bottleneck Suppresses 95-99% of Gradients
A COLM'26 submission finds the LM head suppresses 95-99% of gradient norm in backpropagation, making it both an expressive and optimization bottleneck; the community also reproduced Claude's tokenizer and noted its small vocabulary.
2026-08-15 ~ 2026-08-16 · 3 related posts
- Paper reveals LM Head as gradient bottleneck amid Anthropic vocab size rumors — JFPuget · 2026-08-15
- Paper: LM Head is a Gradient Bottleneck; Anthropic May Have Solved It — zhaoran_wang · 2026-08-15
- Paper: LM Head is a Gradient Bottleneck, losing 95-99% of gradient norms — tokenbender · 2026-08-16