Qwen3.8-Flash-Next Fits on a 48GB MacBook: Pruning + SSD N-gram Table, 39GB RAM

EyalToledano · x · 2026-08-27

Developer Eyal Toledano got Qwen3.8-Flash-Next running on a 48GB MacBook Air: Q4 on MLX uses only 39GB of memory at 28 tok/s decode, and Q8 runs in 50GB. Stock Q4 needs 97GB.

Two stacked tricks:

Results: 288 experts + table on NVMe loads in 6 seconds, 39GB resident, 28 tok/s decode, 600 tok/s prefill, logits identical to the RAM-resident version. Trade-off: HumanEval drops from 93.9% (stock) to 91.5%.

Next up: an MTP drafter and an n-gram lookup drafter for speculative decoding, both built and now benchmarking.

Original post →

More from Infra

Infra channel →