AMD opens Instella-MoE, a fully open 16B MoE model trained on MI300X and MI325X
QuixiAI · x · 2026-07-28
AMD introduced Instella-MoE, its first fully open Mixture-of-Experts language model.
- Scale: 16B total parameters, with 2.8B active parameters per token.
- Training: trained from scratch on AMD Instinct MI300X and MI325X GPUs using AMD-Primus and Miles.
- Release scope: AMD says it is publishing not just weights, but the full end-to-end pipeline: checkpoints from pre-training through RL, training recipes, configurations, data mixtures, and training/inference code.
- Claims: the company says the model reaches state-of-the-art performance among fully open models at its scale.
- Image context: the attached charts compare Instella-MoE against open and dense baselines on base and post-trained/RL benchmarks, showing it near the top of its peer group.
More from Infra
- Half of U.S. servers may sit in tiny rooms, not giant Virginia data centers — aronchick · 2026-07-28
- Morgan Stanley says AI memory prices may peak in Q4 2026 as NAND inventories rise — SumitGup · 2026-07-28
- Iota says Orion-16B is training live on unreliable compute across three continents — markjeffrey · 2026-07-28
- Databricks pitches AI spend governance as token usage outpaces value — matei_zaharia · 2026-07-28
- A company with 2 H200s asks which coding model and vLLM setup can serve 4–10 users — redblood252 · 2026-07-28
- AMD, Intel and chip-equipment demand headline The Circuit’s latest AI infrastructure episode — BenBajarin · 2026-07-28