Dead layers in deep transformers: attention-collapse paper finds a quarter of layers idle at depth 32
ziv_ravid · x · 2026-10-01
A paper on deep attention collapse is drawing researcher attention: Block AttnRes peaks around 20 layers then degrades, and early layers give near-zero weight to their own outputs, effectively just re-reading embeddings. Models of depth 16/24/32 show 2/5/8 dead layers—about a quarter of the network at 32 layers. Inheritune author zivravid links this to reusing early layers of large models and asks whether such layers can be compressed, dropped, or fused, and whether it scales.
Related event: Study Finds Deep Attention Collapse Renders a Quarter of Layers Dead(2 posts)→
More from Research
- Gemini 3.8 Flash (high) hits 84.8% on WeirdML v2, first Flash to beat Gemini 3.1 Pro — teortaxesTex · 2026-10-01
- Researchers steal frontier models' hidden reasoning by replaying encrypted chain-of-thought traces — maksym_andr · 2026-10-01
- Thinking in Geometric Terms: What ReLU, LayerNorm, LoRA and VQ Do to the Data Space — techNmak · 2026-10-01
- Foresight Institute convenes AI for Frontier & Meta Science workshop in San Francisco — juanbenet · 2026-10-01
- London AI x Science Hackathon in October Welcomes Non-Coding Bench Scientists — ersatzben · 2026-10-01
- First superhuman Stratego AI unveiled in Nature paper using RL and test-time compute — zicokolter · 2026-10-01