Limite 1B borrows nanogpt speedrun architecture: NorMuon, MUDD variant and XSA
tensorqt · x · 2026-10-07
Paradigma explains the architectural choices behind Limite 1B: the team heavily leveraged designs that emerged from the nanogpt speedrun challenges, training with NorMuon on matrix-shaped parameters, using a MUDD variant for residual stream mixing, and XSA in the attention layer.
This is part of the Limite 1B - Violetto technical report release thread; see the full report for complete details.
More from Models
- Cohere Labs Debuts Tiny Aya, a Small-Model Family Covering 70+ Languages — Cohere_Labs · 2026-10-07
- Paradigm releases tech report for Violetto, a 1B model trained from scratch for math — tensorqt · 2026-10-07
- Mistral's Le Chonk tops blind human code review among open models, second only to Opus 5 — qtnx_ · 2026-10-07
- DeepSeek V4.1 Flash hits 72.9% on ARC-AGI-2 at $0.13/task, costing 250% more — teortaxesTex · 2026-10-07
- Mistral claims Large 4 is one of the world's strongest AI models for cybersecurity — scaling01 · 2026-10-07
- Mistral Large 4 solves 18 of 19 CTF challenges in official speedrun with tool calls — MistralAI · 2026-10-07