SemiAnalysis: mapping MoE models onto inference hardware

zephyr_z9 · x · 2026-09-22

SemiAnalysis publishes a deep dive on how MoE models map onto inference hardware: MoE changed which tensors activate per token, what must sit close together, and how memory movement, storage and scheduling shape useful throughput. It covers cluster orchestration (NVIDIA Dynamo, Mooncake), inference servers (vLLM, SQlang), the request/turn model, and KV cache mechanics. Paid article by Tanj Bennett.

Original post →

More from Infra

Infra channel →