Qwen3.8-Flash-Next runs with full experts on just 37GB RAM
EyalToledano · x · 2026-08-29
The author shared a breakthrough in MoE inference optimization for Qwen3.8-Flash-Next. By storing 60% of the experts on disk and streaming them to memory on-demand (similar to n-gram streaming), the model achieves 40 tok/s decoding on a M4 Max with only 37GB RAM while using full experts (no pruning).
More from Infra
- Stripe employee calls for agent-friendly internet infrastructure — jeff_weinstein · 2026-08-29
- Cognition hits $900M ARR but may burn $800M on Nvidia servers — thedealdirector · 2026-08-29
- Cisco Partners with Super Micro for Liquid-Cooled AI Racks — Beth_Kindig · 2026-08-29
- Neocloud Lambda Secures $1B Debt to Buy Nvidia Chips — TechCrunch AI · 2026-08-29
- Kimi K3 Full Fine-Tuning Live: GPU Requirements Drop by 40% — ypatil125 · 2026-08-29
- Audit reveals 64 GGUF quants mislabeled across 25 repos — Daxfortuna · 2026-08-29