Speeding Up Inference via SSD Streaming of Experts

gajesh · x · 2026-07-19

Someone proposed a cost-effective approach to run Kimi K3 at 1t/s: when VRAM is near its limit, you can use SSD streaming of experts to load certain experts on-demand from an SSD.

They also noted that Mac machines boast SSD read speeds up to 7GB/s, implying that this setup is practically viable in certain scenarios. The attached image shows Micron's data center SSD page, highlighting the importance of high-bandwidth storage for AI inference and training.

Original post →

More from Infra

Infra channel →