SGLang's SSD Expert Pack runs huge MoE models off an NVMe SSD on one RTX 5090

ying11231 · x · 2026-09-19

The SGLang team and WiCi AI released SSD Expert Pack, letting consumer hardware run models far larger than system RAM: routed experts stay on an NVMe SSD, and the runtime loads only the experts the router selects into a GPU cache.

On one RTX 5090 with 32 GB RAM and a 2 TB SSD:

Original post →

More from Infra

Infra channel →