AMD Enters Open-Source LLM Arena with Instella-MoE-16B
airesearch12 · x · 2026-08-03
AMD has officially entered the open-source LLM space by releasing its own reasoning Mixture-of-Experts (MoE) model, Instella-MoE-16B-A3B-Think.
Model Features:
- 16B total parameters with only 2.8B active parameters per token
- 64 routed experts (6 selected + 2 shared per token)
- Trained from scratch on 7.1T tokens
- Uses an SFT → DPO → RL reasoning pipeline
- Built entirely on AMD Instinct GPUs and ROCm software stack
Performance:
According to AMD-reported evaluations, the model achieves the highest overall average score of 73.22, outperforming peers like OLMo3-7B-Think, Gemma-4-E4B Think, and Qwen3.5-4B. It also takes the top score in AIME25 and LiveCodeBench.
More from Infra
- Qwen3.8-27B Open Weights Coming, Runs Locally on 17GB RAM — danielhanchen · 2026-08-03
- AirLLM Breaks VRAM Barrier: Runs 70B LLMs on a Single 4GB GPU — techNmak · 2026-08-03
- MiniMax H3 Open Weights Hit fal with Out-of-the-Box Inference Optimizations — gorkem · 2026-08-03
- ComfyUI Adds Day 0 Support for MiniMax Video Model, Slashing VRAM by 66% for RTX 3060 — crystal_alpine · 2026-08-03
- US States Move to Repeal Data Center Tax Breaks, Raising AI Infrastructure Costs — pstAsiatech · 2026-08-03
- Handling Offline AI Jobs: Developers Share Best Engineering Practices — cmm324 · 2026-08-03