Dead layers in deep transformers: attention-collapse paper finds a quarter of layers idle at depth 32

ziv_ravid · x · 2026-10-01

A paper on deep attention collapse is drawing researcher attention: Block AttnRes peaks around 20 layers then degrades, and early layers give near-zero weight to their own outputs, effectively just re-reading embeddings. Models of depth 16/24/32 show 2/5/8 dead layers—about a quarter of the network at 32 layers. Inheritune author zivravid links this to reusing early layers of large models and asks whether such layers can be compressed, dropped, or fused, and whether it scales.

Related event: Study Finds Deep Attention Collapse Renders a Quarter of Layers Dead(2 posts)→

Original post →

More from Research

Research channel →