Berkeley's 2.3B hybrid Mamba MoE nears Llama-3.2-3B with under 1% of compute
A UC Berkeley team released Rigel, a 2.3B-parameter MoE hybrid Mamba-2 model (360M active) trained with under 1% of Llama-3.2-3B's compute while approaching its performance.
2026-09-23 ~ 2026-09-24 · 2 related posts
- 2.3B MoE hybrid Mamba model matches Llama-3.2-3B with <1% of its pretraining FLOPs — tri_dao · 2026-09-23
- Berkeley team trains 2.3B MoE Hybrid Mamba matching Llama-3.2-3B with <1% of pretraining FLOPs — berkeley_ai · 2026-09-24