Qwen3.8-Flash-Next with pMLX Engine Runs Big Models on Small Memory

The upcoming Qwen3.8-Flash-Next ships with a custom pMLX engine offering dynamic quantization and NVMe streaming. Benchmarks show it runs with 256k context at 45 tok/s on an M1 Pro, enabling code workloads with just 12GB of memory.

2026-09-01 ~ 2026-09-02 · 3 related posts