Qwen3.8-Flash-Next runs with full experts on just 37GB RAM

EyalToledano · x · 2026-08-29

The author shared a breakthrough in MoE inference optimization for Qwen3.8-Flash-Next. By storing 60% of the experts on disk and streaming them to memory on-demand (similar to n-gram streaming), the model achieves 40 tok/s decoding on a M4 Max with only 37GB RAM while using full experts (no pruning).

Original post →

More from Infra

Infra channel →