Berkeley team trains 2.3B MoE Hybrid Mamba matching Llama-3.2-3B with <1% of pretraining FLOPs

berkeley_ai · x · 2026-09-24

A UC Berkeley academic team unveiled Rigel, a 2.3B-parameter MoE (360M active) Hybrid Mamba-2 model that lands within a few points of Llama-3.2-3B dense using less than 1% of its pretraining FLOPs.

Key points:

A notable demonstration that new architectures can be validated on a shoestring, heterogeneous compute budget.

Related event: Berkeley's 2.3B hybrid Mamba MoE nears Llama-3.2-3B with under 1% of compute(2 posts)→

Original post →

More from Research

Research channel →