Block AttnRes analysis: a quarter of layers dead at depth 32, peak mixing at 20 layers
ziv_ravid · x · 2026-10-01
Commenting on Block AttnRes research, zivravid notes that depth mixing peaks at 20 layers and then degrades. Early layers give near-zero weight to their own outputs, effectively just re-reading embeddings. Models at L=16/24/32 contain 2, 5, and 8 dead layers respectively — at depth 32, roughly a quarter of the network is dead, a significant finding for architecture scaling.
Related event: Study Finds Deep Attention Collapse Renders a Quarter of Layers Dead(2 posts)→
More from Models
- Abliterated, Heretic, and Obliteratus uncensored models now discoverable on Pirate Face — alexcovo_eth · 2026-10-01
- Cohere launches Embed 5: preference-aware retrieval, stronger multilingual, big throughput gains — cohere · 2026-10-01
- Gliner2.5-Decide, a Jev-style model, tops HF trending with little community notice — parepeg · 2026-10-01
- GPT 6.1 Sol claims only 26k tokens left in reasoning while 150k actually remain — Mental_Ice6435 · 2026-10-01
- ChatGPT code projects only offer Sol 5.6 while chats show 6.1 — zeeg · 2026-10-01
- GPT-6 Astra Ultrafast now roughly as fast as Gemini 3.5 Flash Lite — Angaisb_ · 2026-10-01