Qwen3.8-Flash-Next with pMLX Engine Runs Big Models on Small Memory
The upcoming Qwen3.8-Flash-Next ships with a custom pMLX engine offering dynamic quantization and NVMe streaming. Benchmarks show it runs with 256k context at 45 tok/s on an M1 Pro, enabling code workloads with just 12GB of memory.
2026-09-01 ~ 2026-09-02 · 3 related posts
- Qwen3.8 pMLX Engine: Runs on 12GB RAM at 10 tok/s — EyalToledano · 2026-09-01
- New Qwen 3.8 MLX Engine Features Dynamic Quantization — MgkMshrmBrkfst · 2026-09-01
- M1 Pro Benchmarks: Qwen 3.8 Hits 45 tok/s at 256k Context — EyalToledano · 2026-09-02