Paradigm borrows nanogpt speedrun tricks: NorMuon, MUDD variant and XSA in its training stack

PMinervini · x · 2026-10-07

Paradigm revealed it has strongly leveraged and adapted architectural choices from the nanogpt speedrun challenges: training matrix-shaped parameters with NorMuon, using a MUDD variant for residual stream mixing, and XSA in the attention layer. The retweeter notes that speedruns have become a treasure trove for architectural innovation — evidence that community-driven speedrun techniques are being adopted in real industrial training projects.

Related event: Paradigm Training Stack Borrows Heavily from nanogpt speedrun(3 posts)→

Original post →

More from Research

Research channel →