Paradigm borrows nanogpt speedrun tricks: NorMuon, MUDD variant and XSA in its training stack
PMinervini · x · 2026-10-07
Paradigm revealed it has strongly leveraged and adapted architectural choices from the nanogpt speedrun challenges: training matrix-shaped parameters with NorMuon, using a MUDD variant for residual stream mixing, and XSA in the attention layer. The retweeter notes that speedruns have become a treasure trove for architectural innovation — evidence that community-driven speedrun techniques are being adopted in real industrial training projects.
Related event: Paradigm Training Stack Borrows Heavily from nanogpt speedrun(3 posts)→
More from Research
- Lampinen vs Bowers debate: do computers amplify minds or compute with symbols? — AndrewLampinen · 2026-10-07
- Apple paper: a single well-prompted agent with shell beats multi-agent ML harnesses, 62.5% vs 47.1% Kaggle medal rate — rohanpaul_ai · 2026-10-07
- Chan Zuckerberg Biohub Hosts Inaugural AIxBio Event in San Francisco — lucapinello · 2026-10-07
- FIRM needs only 10 NFEs and 0.24s per image on CelebA deblurring, beating baselines — prof_kamilov · 2026-10-07
- First-token logprob gating fixes SKIP recall in a local real-time medical scribe, with zero false skips — r-chop14 · 2026-10-07
- Virginia Tech's Hybrid Latent Attention boosts looped LLM GPU throughput up to 8.8x with minimal accuracy loss — rohanpaul_ai · 2026-10-07