DepthBench paper finds Pre-LN variants hit a depth wall, comparing 10 residual designs
teortaxesTex · x · 2026-09-30
Shiwei Liu's team introduces a new scaling axis — effective computational depth — and releases DepthBench, an apples-to-apples comparison of 10 residual connection designs (AttnRes, mHC, KEEL, LNS, MoDA, etc.).
Setup: fix model size, data and training recipe, and vary only the width–depth ratio from shallow-wide to deep-narrow. Key findings:
- Pre-LN and most of its norm/scaling variants hit a wall: pushing more capacity into depth flattens then reverses gains, with optima at large aspect ratios of 42.7–76.0.
- In an experimental 400M model, mHC saturates at 24 layers, but in the regime targeted by very wide, shallow models like DSV4.1-Flash, it is likely at least no worse than AttnRes.
The author calls for internal lab ablations to validate the results at scale.
Related event: DepthBench: Residual Connection Design Caps Effective Computational Depth(2 posts)→
More from Models
- Sentdex: GLM 5.3 Puts a Frontier Model on Your Home Machine — Stop Paying Companies to Lecture You — Sentdex · 2026-09-30
- LLM Chain-of-Thought Contains Surprising Emotionally Expressive Language—and It May Be Functional — xuanalogue · 2026-09-30
- When the CoT Says 'Responding with Honest Feedback,' That's When You Worry — TheZvi · 2026-09-30
- Anthropic 'Drops a Banger Gift' for Claude Users, Says Popular AI Blogger — eyishazyer · 2026-09-30
- GPT 6.1 Sol cuts cached pricing 50% while Anthropic holds back models over safety — oran_ge · 2026-09-30
- OpenRouter data: token usage exploding, some open-weight models see 10x spend since January — AccBalanced · 2026-09-30