oMLX 0.6.3 Released: Qwen/GLM Flash Support, 160% Faster Decode

HankYeomans · x · 2026-08-28

oMLX 0.6.3 adds first-class support for Qwen3.8-Flash-Next and GLM-5.3-Flash with oQe quantization. Qwen3.8-Flash-Next combines Lightning MTP with SSD offloading, boosting 4K decode from 22.2 to 58.1 tok/s on M3 Ultra with 96.8% acceptance. The experimental GPU+ANE+CPU path also sees improved stability and long-prompt quality.

Original post →

More from Infra

Infra channel →