Limite 1B borrows nanogpt speedrun architecture: NorMuon, MUDD variant and XSA

tensorqt · x · 2026-10-07

Paradigma explains the architectural choices behind Limite 1B: the team heavily leveraged designs that emerged from the nanogpt speedrun challenges, training with NorMuon on matrix-shaped parameters, using a MUDD variant for residual stream mixing, and XSA in the attention layer.

This is part of the Limite 1B - Violetto technical report release thread; see the full report for complete details.

Original post →

More from Models

Models channel →