Multi-Head Attention Residuals Enhance Transformer Depth Routing

heghbalz · x · 2026-08-02

Traditional Transformers propagate information across depth using a single additive residual stream, forcing all feature subspaces to share one routing query. This compromise degrades performance as model width grows.

The paper Multi-Head Attention Residuals (MHAR) introduces a new architecture to fix this:

Related event: MHAR Improves Transformer Residuals to Reduce Validation Loss(2 posts)→

Original post →

More from Research

Research channel →