xHC: A New Architecture Breaking Transformer Residual Stream Scaling Bottlenecks
rednote-hilab · hf · 2026-07-20
Existing hyper-connection methods are typically limited when scaling beyond 4 parallel residual streams due to diminishing returns and soaring training costs. The study notes this is mainly constrained by insufficient write-back information and the cubic growth of residual mixing costs.
To address this, the research introduces xHC (Expanded Hyper-Connections), the first method to achieve effective scaling for N>4. Its core designs include:
- Temporal Feature Enhancement: Provides richer write-back information.
- Sparse Residual Stream Architecture: Updates only k=4 out of N=16 streams while maintaining dense access to the global state.
Experiments on 18B and 28B MoE models show that xHC delivers significant downstream performance improvements with moderate training overhead. Furthermore, the proposed xHC-Flash technique effectively controls memory bandwidth traffic, reducing overhead to a level comparable to mHC with N=4, making large-scale residual stream scaling genuinely practical.
Related event: xHC Architecture Breaks Transformer Residual Stream Limits(2 posts)→
More from Research
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Jeff Heaton's Intro to the Math of Neural Networks eBook Is Free to Download — blaizedsouza · 2026-09-11
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11