Qwen3.8-Flash-Next Fits on a 48GB MacBook: Pruning + SSD N-gram Table, 39GB RAM
EyalToledano · x · 2026-08-27
Developer Eyal Toledano got Qwen3.8-Flash-Next running on a 48GB MacBook Air: Q4 on MLX uses only 39GB of memory at 28 tok/s decode, and Q8 runs in 50GB. Stock Q4 needs 97GB.
Two stacked tricks:
- REAP pruning cut the expert pool from 68GB to 35GB (288 experts kept);
- The 51B n-gram table now lives on SSD instead of RAM — his patch memmaps the shards, reads 100 bytes per lookup straight off disk, dequantizes on CPU, verified bit-exact against the GPU path.
Results: 288 experts + table on NVMe loads in 6 seconds, 39GB resident, 28 tok/s decode, 600 tok/s prefill, logits identical to the RAM-resident version. Trade-off: HumanEval drops from 93.9% (stock) to 91.5%.
Next up: an MTP drafter and an n-gram lookup drafter for speculative decoding, both built and now benchmarking.
More from Infra
- Running frontier models locally on Mac is brutal but possible work — dscape · 2026-08-27
- Chinese models iterate 2x faster, cutting inference costs by 80% in 8 weeks — bindureddy · 2026-08-27
- Llama.cpp on ROCm Crashes Constantly on Radeon 780M—One Env Var Fixes It — MaximusSenior · 2026-08-27
- Kioxia and Sandisk to Invest Over $31 Billion in Japan Amid AI Boom — badumtsssst · 2026-08-27
- yacine gets real-time stereo depth latency down to imperceptible levels — yacineMTB · 2026-08-27
- Dev builds a native iOS app to set up and cluster DGX Sparks in a few taps — HankYeomans · 2026-08-27