Paradigm adapts nanogpt speedrun tricks with NorMuon, MUDD variant and XSA to balance sample efficiency and training speed
tensorqt · x · 2026-10-07
- The Paradigm team tackles a core tension in pretraining: limited high-quality data demands sample efficiency, but better sample efficiency typically means longer training time. They set out to find a sweet spot balancing both speed and efficiency.
- Their approach heavily leverages and adapts architectural choices that emerged from the nanogpt speedrun challenges:
- Trains matrix-shaped parameters with NorMuon, a Muon optimizer variant
- Uses a MUDD variant for residual stream mixing
- Employs XSA in the attention layer
- The work is an empirical exploration of training-method design that could inform efficiency-focused pretraining research.
Related event: Paradigm Training Stack Borrows Heavily from nanogpt speedrun(3 posts)→
More from Research
- Fixed token codes suffice: 1.7B LM trains without a trainable input embedding table — A. Bochkov · 2026-10-07
- EmbeddingGemma 2 hands-on: 740M multimodal embeddings for search and RAG, runnable on a free T4 — Prompt Engineering · 2026-10-07
- Isomorphic, DeepMind and Meta join DOE-NIH partnership to build an AI model of the cell — snikolov · 2026-10-07
- Bi-manual mobile UMI demo unlocked for robot manipulation data collection — neurosp1ke · 2026-10-07
- Researcher presents Meta-Harness and Combee at COLM 2026, seeks industry roles — Kangwook_Lee · 2026-10-07
- Lampinen: great cultural insights rarely come from a single brain, unlike LLM analogy — AndrewLampinen · 2026-10-07