Paper: LM Head Bottleneck Suppresses 95-99% of Gradients

A COLM'26 submission finds the LM head suppresses 95-99% of gradient norm in backpropagation, making it both an expressive and optimization bottleneck; the community also reproduced Claude's tokenizer and noted its small vocabulary.

2026-08-15 ~ 2026-08-16 · 3 related posts