xHC: A New Architecture Breaking Transformer Residual Stream Scaling Bottlenecks
rednote-hilab · hf · 2026-07-20
Existing hyper-connection methods are typically limited when scaling beyond 4 parallel residual streams due to diminishing returns and soaring training costs. The study notes this is mainly constrained by insufficient write-back information and the cubic growth of residual mixing costs.
To address this, the research introduces xHC (Expanded Hyper-Connections), the first method to achieve effective scaling for N>4. Its core designs include:
- Temporal Feature Enhancement: Provides richer write-back information.
- Sparse Residual Stream Architecture: Updates only k=4 out of N=16 streams while maintaining dense access to the global state.
Experiments on 18B and 28B MoE models show that xHC delivers significant downstream performance improvements with moderate training overhead. Furthermore, the proposed xHC-Flash technique effectively controls memory bandwidth traffic, reducing overhead to a level comparable to mHC with N=4, making large-scale residual stream scaling genuinely practical.
Related event: xHC Architecture Breaks Transformer Residual Stream Limits(2 posts)→
More from Research
- Gigatoken: Open-Source Tokenizer Claiming 100x Speedup Over Tiktoken — Thrumpwart · 2026-07-22
- Can Kimi or GLM replicate recent closed-model math and cyber breakthroughs offline? — Unusual_Guidance2095 · 2026-07-22
- Kimi K3 is said to match Fable in a new SOTA comparison — piotrgrabowski · 2026-07-22
- Krea 2 LoKr likeness guide says 750 steps is usually enough for near-perfect face training — LilBrownBebeShoes · 2026-07-22
- PoLar: Dynamically Skipping or Looping LLM Layers for Efficient Inference — ttkciar · 2026-07-22
- Stanford Paper Examines the Institutional Context of AI Benchmarks for Consumers and Regulators — chrmanning · 2026-07-22