Would 100+ layers tip the scales between attnres and full attention residual?
stochasticchasm · x · 2026-08-28
A researcher raises a question about attention residual (attnres) vs. full attention residual: their benchmark results are already very close, and at more extreme depth (100+ layers) the rankings might flip, since attnres can always pull in full-magnitude residuals in early layers. In a follow-up reply, they note two design details: the 2x factor in the residual write-back op seems tied to init dynamics — sigmoid(0)=0.5 with zero-centered init, so a 2×sigmoid restores normal-scale writeback — and the \overline{R} notation appears unused anywhere, despite group-RMSNorm being called out as important in the prior section.
Related event: Residual Write-Back Design and AttnRes Depth Debate Draw Attention(2 posts)→
More from Research
- Michigan Robotics Rounds Up Its Papers and Workshops for IROS 2026 — doctorBobG · 2026-08-28
- Mouse Brain Connectome Cost Drops to $100M; Human at $1B — juanbenet · 2026-08-28
- Testing Muon Optimizer: Smoother Gradients and Stable Residual Maxima — stochasticchasm · 2026-08-28
- First open model adopts per-head orthogonalization following Kimi and GLM — stochasticchasm · 2026-08-28
- Integrity Bench: A New Benchmark to Measure Model Overconfidence — Acne_Discord · 2026-08-28
- Discussion on Classic Moonlight Scaling and Polar Express Orthogonalization — stochasticchasm · 2026-08-28