DepthBench compares 10 architectures to find which residual tweaks actually buy computational depth
SonglinYang4 · x · 2026-09-30
Frontier labs are redesigning the residual stream to beat the curse of depth: Kimi K3 uses AttnRes, DeepSeek V4 uses mHC, ByteDance proposed HC, plus LNS, KEEL, MoDA. But each was validated with its own training budget and codebase, so results aren't comparable. DepthBench (arXiv:2609.32534) is a controlled benchmark that fixes model size and pre-training recipe while varying the width–depth ratio across 10 architectures. Key findings: depth allocation gains are strongly architecture-dependent; standard Pre-LN and most norm/scaling variants offer little benefit and can degrade as models get deeper and narrower; HC and Full AttnRes keep improving even at extreme deep-narrow shapes, with gains carrying from pre-training loss to downstream performance.
More from Models
- Sam Altman asks how OpenAI should charge for Codex, $500/mo plan speculation spreads — chaumian · 2026-09-30
- ChatGPT Pro's $200 plan reportedly includes 62,500 Codex credits expiring Dec 31 — chaumian · 2026-09-30
- User calculates 62,500 credits ≈ $2,500 of GPT-6.1 API usage, calling the new plan a big cut — chaumian · 2026-09-30
- Reddit user reports surprise 62,500 credit grant, about 15x their monthly plan allowance — IronDarbe · 2026-09-30
- Users question why ChatGPT chat mode still misses the newest GPT models — koltregaskes · 2026-09-30
- User receives 62,500 credits worth ~$2,500, equal to 12.5 months of the $200 Pro plan — kimmonismus · 2026-09-30