Paper reveals LM Head as gradient bottleneck amid Anthropic vocab size rumors
JFPuget · x · 2026-08-15
Addressing rumors about Anthropic's model vocabulary size, this post cites the paper "Lost in Backpropagation: The LM Head is a Gradient Bottleneck". The paper argues that projecting output features to the vocabulary dimension creates not just an expressivity bottleneck, but an optimization bottleneck. Backpropagating V-dimensional gradients through a rank-D linear layer causes unavoidable compression, altering training feedback for most parameters. Empirical measurements show 95-99% of gradient norm is suppressed by the output layer, leading to suboptimal update directions.
More from Models
- Qwen 3.8 27B Beats Claude Opus 4.6 in Three.js Coding Test for Free — testingcatalog · 2026-08-15
- SemiAnalysis: GLM-5.3 beats US open models via post-training gains — teortaxesTex · 2026-08-15
- Grok 4.6 leads spatial biology benchmark but fails biosecurity tests — kenbwork · 2026-08-15
- Gemini underrated: outperforms big models in retrieval and math — rickasaurus · 2026-08-15
- Ex-OpenAI researcher points out issues in Grok 4.6 system card — Miles_Brundage · 2026-08-15
- Zhipu GLM-5.3 Rivals Frontier Models: How Chinese Labs Keep Pace — Interconnects (Nathan Lambert) · 2026-08-15