Paper: LM Head is a Gradient Bottleneck; Anthropic May Have Solved It
zhaoran_wang · x · 2026-08-15
A new paper identifies the LM Head as a gradient bottleneck in LLMs, suppressing 95-99% of gradient norms and harming training efficiency. Concurrently, a reproduction of Claude's tokenizer revealed a surprisingly small vocabulary of 15k entries, leading to speculation that Anthropic has mitigated this bottleneck by using a reduced vocabulary size.
Related event: Paper: LM Head Bottleneck Suppresses 95-99% of Gradients(3 posts)→
More from Models
- Anthropic details how Claude’s new watermarks work: mechanism and resistance to editing — TechCrunch AI · 2026-08-16
- Qwen3.8-27B on 16GB VRAM: KV Cache Quantization Cliff from q4_0 to q4_1 — Unnamed-3891 · 2026-08-16
- Benchmark: Qwen3.8-27B doubles coding ability vs Qwen3.6-27B — poppear · 2026-08-16
- Qwen 3.8 27B Beats Codex in Coding Benchmarks: Wins 8/13, Costs 1/3 — tokenbender · 2026-08-16
- DeepSeek Accused of Grey Testing for High Scores, Opus 5 Output Quality Questioned — teortaxesTex · 2026-08-16
- Study: Cross-Version Transfer of Qwen Interpretability Lenses — imstilllearningthis · 2026-08-16