Paper: LM Head is a Gradient Bottleneck, losing 95-99% of gradient norms

tokenbender · x · 2026-08-16

A paper submitted to COLM'26 reveals that the final layer of Language Models (LMs) is not just an expressivity bottleneck but also an optimization bottleneck. Projecting D-dimensional outputs to V-dimensional logits (where D << V) causes unavoidable gradient compression during backpropagation.

Key Findings:

Related event: Paper: LM Head Bottleneck Suppresses 95-99% of Gradients(3 posts)→

Original post →

More from Research

Research channel →