AMD opens Instella-MoE, a fully open 16B MoE model trained on MI300X and MI325X
QuixiAI · x · 2026-07-28
AMD introduced Instella-MoE, its first fully open Mixture-of-Experts language model.
- Scale: 16B total parameters, with 2.8B active parameters per token.
- Training: trained from scratch on AMD Instinct MI300X and MI325X GPUs using AMD-Primus and Miles.
- Release scope: AMD says it is publishing not just weights, but the full end-to-end pipeline: checkpoints from pre-training through RL, training recipes, configurations, data mixtures, and training/inference code.
- Claims: the company says the model reaches state-of-the-art performance among fully open models at its scale.
- Image context: the attached charts compare Instella-MoE against open and dense baselines on base and post-trained/RL benchmarks, showing it near the top of its peer group.
Related event: AMD Releases Fully Open-Source MoE Model Instella(2 posts)→
More from Infra
- fmgo: call Apple's on-device Foundation Models from Go with no CGO and no Swift — Super_Run_8466 · 2026-09-23
- Huawei unveils Peerium architecture: nested BSP unifies million processors into one computer — Dr_Singularity · 2026-09-23
- Grok explains why DeepSeek picked DualPipe + ZeRO-1 over ZeRO-3 on 2048 H800s — TheZachMueller · 2026-09-23
- AI costs fall 47% per quarter, 4x faster than DNA sequencing: Epoch AI — daveholtz · 2026-09-23
- M5 Ultra LLM test: 4x faster prompt processing, but double the power draw — DigitalguyCH · 2026-09-23
- $500 of Dell OptiPlexes become a diskless netboot lab where AI agents can't brick the hardware — colinmcnamara · 2026-09-23