New theory shows how to scale residual network updates when layer weights are correlated
burkov · x · 2026-09-06
Andrey Burkov breaks down a recent paper on very deep residual networks: each layer's update must be carefully scaled — too small and the network approaches the identity map, too large and hidden states become unstable. The paper develops scaling theory for initializations with long-range cross-layer correlations, showing the correct scaling depends on how fast correlations decay and how initial weights are generated — making these meaningful initialization hyperparameters. It also characterizes what ultra-deep networks converge to beyond the previously known ODE and SDE regimes.
More from Research
- New paper fixes contrastive RL blind spot by re-weighting InfoNCE with 1-bit failure signal — kastnerkyle · 2026-09-06
- NBER: AI investment boosts firm productivity since 2018 by building organization capital — TaniaBabina · 2026-09-06
- Thomas Kipf: intelligence is minimizing the generator-verifier gap — tkipf · 2026-09-06
- Models over-edit code written by other models; CROCODIL training framework fixes it — omarsar0 · 2026-09-06
- Bare coding agent hits 78% on R2R-CE navigation with zero training, no mapping or memory — jiqizhixin · 2026-09-06
- Slow Clinical Trials Break AI's Feedback Loop in Drug Development — clarejtbirch · 2026-09-06