MoE inference: active parameters matter more than total size
giffmana · x · 2026-07-20
A reply arguing that with MoE, SSD streaming makes local LLM inference feasible even at batch size 1, so the total parameter count matters less than the number of active parameters.
The key claim is that reducing active parameters is always a win, and the remaining tradeoff is mainly that matmul efficiency can become poor in this setup.
Related event: MoE and SSD Streaming Redefine Inference Parameters(3 posts)→
More from Infra
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- DeepSeek-V4-Flash tops out at 770 tok/s on one B300 in a vLLM batch test — Moreh · 2026-07-22
- NVIDIA starts shipping 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-22
- Apple publishes SOC 3 audit reports for Private Cloud Compute — throwfaraway4 · 2026-07-22
- Reddit GPU renters say existing platforms only give you two of three: code, recovery, fair billing — legendpizzasenpai · 2026-07-22
- The Sandboxing Manifesto: Secure Execution Environments for Agents — spirosoik · 2026-07-22