Running MoE via SSD streaming on a 64GB Mac mini: GPU idles 27% of decode waiting on experts
turtleninja99 · reddit · 2026-10-05
The author runs Qwen Flash Next q4 on a 64GB Mac mini (M5), too large to fit fully in memory, using a hybrid scheme: hot experts stay cached, the rest stream from SSD.
- Current performance: 17.5 tok/s decode, 390 tok/s prefill
- Bottleneck: the GPU waits on SSD expert streaming for 27% of decode time
- Optimizations already done: carousel buffering for prompt processing (loads faster than GPU consumes), MTP, lookahead prefetch of next-layer experts (72% fetch accuracy), 75% hot-cache hit rate
- Ideas being explored: expanding lookahead prefetch, a separate staging buffer for guessed experts to avoid evicting hot ones
Open-sourced at Flash-next-ssd; author seeks ideas to raise GPU utilization.
Related event: Qwen Flash Next Runs Locally on 64GB Mac mini M5(2 posts)→
More from Infra
- Use Magpie CLI to check quotas and route sub-agents by urgency to save tokens — lxfater · 2026-10-05
- Aleph Alpha details scaling a 30B MoE pre-training to 512 B200 GPUs at 35.3% MFU — bodonoghue85 · 2026-10-05
- The next AI bottleneck isn't GPUs — it's a gigawatt connected to the grid — ingliguori · 2026-10-05
- AI Is Cheap to Use, Expensive to Provide: Investor Says Compute Needs Real Spot Prices — Kyrannio · 2026-10-05
- Debugging OpenCode stalls: Qwen3.6 hybrid memory forces llama-server full prompt re-processing — MysteriousInterest32 · 2026-10-05
- AI systems may consume 100x the power they need, warns infrastructure analyst — DavidLinthicum · 2026-10-05