Berkeley team trains 2.3B MoE Hybrid Mamba matching Llama-3.2-3B with <1% of pretraining FLOPs
berkeley_ai · x · 2026-09-24
A UC Berkeley academic team unveiled Rigel, a 2.3B-parameter MoE (360M active) Hybrid Mamba-2 model that lands within a few points of Llama-3.2-3B dense using less than 1% of its pretraining FLOPs.
Key points:
- No dedicated cluster: the run hopped between H100s, A100s, V100s (yes, V100s) and TPU v5p/v6e on a single codebase.
- The team argues academia can compensate for scarce compute with creative students training novel architectures on whatever hardware they can find.
- They are seeking resources to scale the research to the next training rung.
A notable demonstration that new architectures can be validated on a shoestring, heterogeneous compute budget.
More from Research
- BFL shows FLUX 3 Action flying a simulated drone, hinting at broader uses — bfl_ai · 2026-09-24
- BFL open-sources 7B world action model FLUX 3 Action, tops RoboLab 6.1pp ahead — bfl_ai · 2026-09-24
- IKEA Assembly Benchmark: Top Model Score Jumped From 28% to 80% in 10 Months — emollick · 2026-09-24
- Juan Perdomo, researcher on performative prediction, joins NYU as assistant professor — thegautamkamath · 2026-09-24
- Claude Opus 5.5 and GPT-6 Sol land on Biomni Lab research platform — KexinHuang5 · 2026-09-24
- JevK5 open decision model ranks #2 of 76 on JevBench, outputs calibrated probabilities in ~13ms — airesearch12 · 2026-09-24