Running MoE via SSD streaming on a 64GB Mac mini: GPU idles 27% of decode waiting on experts

turtleninja99 · reddit · 2026-10-05

The author runs Qwen Flash Next q4 on a 64GB Mac mini (M5), too large to fit fully in memory, using a hybrid scheme: hot experts stay cached, the rest stream from SSD.

Open-sourced at Flash-next-ssd; author seeks ideas to raise GPU utilization.

Related event: Qwen Flash Next Runs Locally on 64GB Mac mini M5(2 posts)→

Original post →

More from Infra

Infra channel →