2.3B MoE hybrid Mamba model matches Llama-3.2-3B with <1% of its pretraining FLOPs
tri_dao · x · 2026-09-23
A developer pretrained Rigel, a 2.3B MoE (360M active) hybrid Mamba-2 model that lands within a few points of Llama-3.2-3B using less than 1% of its pretraining FLOPs.
- No dedicated cluster: the single run hopped between H100s, A100s, V100s and TPU v5p/v6e on one codebase.
- Tri Dao amplified the thread, praising the effort of getting pretraining to work across three generations of Nvidia GPUs and two generations of TPUs, calling it a very good model for its size.
More from Infra
- fmgo: call Apple's on-device Foundation Models from Go with no CGO and no Swift — Super_Run_8466 · 2026-09-23
- Huawei unveils Peerium architecture: nested BSP unifies million processors into one computer — Dr_Singularity · 2026-09-23
- Grok explains why DeepSeek picked DualPipe + ZeRO-1 over ZeRO-3 on 2048 H800s — TheZachMueller · 2026-09-23
- AI costs fall 47% per quarter, 4x faster than DNA sequencing: Epoch AI — daveholtz · 2026-09-23
- M5 Ultra LLM test: 4x faster prompt processing, but double the power draw — DigitalguyCH · 2026-09-23
- $500 of Dell OptiPlexes become a diskless netboot lab where AI agents can't brick the hardware — colinmcnamara · 2026-09-23