MoE inference: active parameters matter more than total size

giffmana · x · 2026-07-20

A reply arguing that with MoE, SSD streaming makes local LLM inference feasible even at batch size 1, so the total parameter count matters less than the number of active parameters.

The key claim is that reducing active parameters is always a win, and the remaining tradeoff is mainly that matmul efficiency can become poor in this setup.

Related event: MoE and SSD Streaming Redefine Inference Parameters(3 posts)→

Original post →

More from Infra

Infra channel →