DepthBench compares 10 architectures to find which residual tweaks actually buy computational depth

SonglinYang4 · x · 2026-09-30

Frontier labs are redesigning the residual stream to beat the curse of depth: Kimi K3 uses AttnRes, DeepSeek V4 uses mHC, ByteDance proposed HC, plus LNS, KEEL, MoDA. But each was validated with its own training budget and codebase, so results aren't comparable. DepthBench (arXiv:2609.32534) is a controlled benchmark that fixes model size and pre-training recipe while varying the width–depth ratio across 10 architectures. Key findings: depth allocation gains are strongly architecture-dependent; standard Pre-LN and most norm/scaling variants offer little benefit and can degrade as models get deeper and narrower; HC and Full AttnRes keep improving even at extreme deep-narrow shapes, with gains carrying from pre-training loss to downstream performance.

Related event: DepthBench: Residual Connection Design Decides Whether Depth Becomes the Next Scaling Axis(5 posts)→

Original post →

More from Models

Models channel →