Qwen3.8 pMLX Engine: Runs on 12GB RAM at 10 tok/s

EyalToledano · x · 2026-09-01

Qwen3.8-Flash-Next is launching with a custom pMLX engine supporting tiered models and dynamic quantization (bf16/q8/q4/q3). Features include on-the-fly expert pruning, adjustable n-gram streaming, and NVME offloading. Benchmarks show code editing at 58 tok/s sustained on 32k context (Q4), and a low-memory config running at 10 tok/s on just 12GB RAM (Q3).

Original post →

More from Infra

Infra channel →