SGLang's SSD Expert Pack runs huge MoE models off an NVMe SSD on one RTX 5090
ying11231 · x · 2026-09-19
The SGLang team and WiCi AI released SSD Expert Pack, letting consumer hardware run models far larger than system RAM: routed experts stay on an NVMe SSD, and the runtime loads only the experts the router selects into a GPU cache.
On one RTX 5090 with 32 GB RAM and a 2 TB SSD:
- DeepSeek-V4-Flash MXFP4: 1.85–1.99 tokens/sec decode
- Kimi-K3 community Q2K (text-only): 0.29 tokens/sec decode
More from Infra
- MiMo RL livestream: infra restarts waste 33% of compute, over $400K lost — dustinvtran · 2026-09-19
- Why Jev-class models could become a near-free judgment primitive running on-device — signulll · 2026-09-19
- Pedro Domingos: Opposing Data Centers Means Keeping Your Country Stupid — pmddomingos · 2026-09-19
- Pedro Domingos: OpenAI and Anthropic's real moat is their massive secured compute — pmddomingos · 2026-09-19
- Fully Local, Private Computer-Use AI Arrives — Runs on a 3060 Ti 12GB — WolframRvnwlf · 2026-09-19
- AWS ships 13 SageMaker inference updates in 2026, cutting cold-start latency up to 65% — AWS ML Blog · 2026-09-19