Paper: LM Head is a Gradient Bottleneck; Anthropic May Have Solved It

zhaoran_wang · x · 2026-08-15

A new paper identifies the LM Head as a gradient bottleneck in LLMs, suppressing 95-99% of gradient norms and harming training efficiency. Concurrently, a reproduction of Claude's tokenizer revealed a surprisingly small vocabulary of 15k entries, leading to speculation that Anthropic has mitigated this bottleneck by using a reduced vocabulary size.

Related event: Paper: LM Head Bottleneck Suppresses 95-99% of Gradients(3 posts)→

Original post →

More from Models

Models channel →