Representation Collapse and Fix in Pre-Norm

SeunghyunSEO7 · x · 2026-07-10

A reply mentions that MAI observed a similar issue and mitigated it by initializing the output projection of the attention module to 0.0, allowing the MLP to learn first before the representation gradually evolves. The preceding context also notes that representation collapse in pre-norm architectures might harm the load balancing of MoE, with related improvement ideas including depth scaling and muon.

Related event: Pre-norm Architecture May Disrupt MoE Load Balancing(2 posts)→

Original post →

More from Research

Research channel →