DepthBench: residual connection design decides whether depth is a real scaling axis
Intelligent-Systems · hf · 2026-09-29
DepthBench is a controlled benchmark varying width–depth aspect ratio at fixed model size and training recipe across 10 architectures to measure effective computational depth.
- Standard Pre-LN and most norm/scaling variants show little benefit and can degrade as models go deeper and narrower
- HC and Full AttnRes improve consistently even at extreme deep-narrow shapes, extending to domain-specific performance
- Layer-level analyses link their gains to better utilization of additional layers
Conclusion: residual connection design is a key determinant of whether architectural depth translates into effective computation.
More from Research
- ColNanoVDR distills multi-vector visual document retrieval without documents, keeping 95% NDCG@5 at 149M params — nanovdr · 2026-09-29
- NUS rethinks DiT residual connectivity: 1.73x fewer training iterations, 1.39 FID — NationalUniversityofSingapore · 2026-09-29
- Imprint Reader Decodes Weight Updates into Natural Language, Enables Targeted Edits — Guanxu Chen · 2026-09-29
- SJTU's GeoVerse Synthesizes World-Consistent Novel Views in Geometric Latent Space — SJTU · 2026-09-29
- Tencent Hunyuan Maps Scaling Laws for Encoder-Free Multimodal Pretraining — Tencent-Hunyuan · 2026-09-29
- When Do Model Internals Help? Benchmarking Representation Engineering for LLM Safety — Tianyi Guan · 2026-09-29