'Depth Delusion' paper: Transformers should scale width 2.8x faster than depth
xuanalogue · x · 2026-09-04
A new arXiv paper proposes architecture-conditioned scaling laws: optimal depth scales as DC^0.12 while optimal width scales as WC^0.34 — width should grow 2.8x faster. Beyond a critical depth DcritW^0.44, adding layers increases loss despite more parameters. Validated across 30 transformer architectures (17M–7B params, R²=0.922); at 7B scale, a 64-layer 6.38B model underperforms a 32-layer 6.86B model by 0.12 nats. Deeper isn't better at production scale.
More from Research
- DeepMind's "LLM can't jump" argument: induction and deduction can't produce scientific revolutions — 0xsachi · 2026-09-04
- Free 39-Episode Control Bootcamp: The Control Theory That Runs Real Robots — lukas_m_ziegler · 2026-09-04
- UniReps Workshop in Paris opens call for papers on unified representations, due Oct 4 — ClementineDomi6 · 2026-09-04
- Utopia: open-source bitemporal knowledge graph gives RAG agents a memory of change — Shruti_0810 · 2026-09-04
- Emotion is an optimizer's control plane, not an irrational advisor — mimi10v3 · 2026-09-04
- New f-loss Cures Spectral Bias in Pixel-Space Flow Matching, Speeding Convergence — serrjoa · 2026-09-04