Block AttnRes analysis: a quarter of layers dead at depth 32, peak mixing at 20 layers

ziv_ravid · x · 2026-10-01

Commenting on Block AttnRes research, zivravid notes that depth mixing peaks at 20 layers and then degrades. Early layers give near-zero weight to their own outputs, effectively just re-reading embeddings. Models at L=16/24/32 contain 2, 5, and 8 dead layers respectively — at depth 32, roughly a quarter of the network is dead, a significant finding for architecture scaling.

Related event: Study Finds Deep Attention Collapse Renders a Quarter of Layers Dead(2 posts)→

Original post →

More from Models

Models channel →