SemiAnalysis: mapping MoE models onto inference hardware
zephyr_z9 · x · 2026-09-22
SemiAnalysis publishes a deep dive on how MoE models map onto inference hardware: MoE changed which tensors activate per token, what must sit close together, and how memory movement, storage and scheduling shape useful throughput. It covers cluster orchestration (NVIDIA Dynamo, Mooncake), inference servers (vLLM, SQlang), the request/turn model, and KV cache mechanics. Paid article by Tanj Bennett.
More from Infra
- One of the biggest AI launches ever stayed up under crazy load, out-uptimeing Anthropic — hardimanjames · 2026-09-22
- CPU:GPU Ratios and the Race to the Scale Up Domain: Agentic AI Is Reshaping Datacenter CPU Demand — BenBajarin · 2026-09-22
- AMD details EPYC Venice: 2.24x SPECrate lead over NVIDIA Vera, 256-core flagship — ryanshrout · 2026-09-22
- python-build-standalone enables full LTO for CPython 3.12+, modestly boosting runtime — charliermarsh · 2026-09-22
- Measured trade-offs of three REAP-pruned Qwen3.8-Flash-Next MLX builds on Apple Silicon — MensaProdigy · 2026-09-22
- Dev claims further-optimized DeepSeek V4 NVFP4 uses 190GB of 192GB VRAM — HankYeomans · 2026-09-22