MoE inference: active parameters matter more than total size
giffmana · x · 2026-07-20
A reply arguing that with MoE, SSD streaming makes local LLM inference feasible even at batch size 1, so the total parameter count matters less than the number of active parameters.
The key claim is that reducing active parameters is always a win, and the remaining tradeoff is mainly that matmul efficiency can become poor in this setup.
Related event: MoE and SSD Streaming Redefine Inference Parameters(3 posts)→
More from Infra
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11