oMLX 0.6.3 Released: Qwen/GLM Flash Support, 160% Faster Decode
HankYeomans · x · 2026-08-28
oMLX 0.6.3 adds first-class support for Qwen3.8-Flash-Next and GLM-5.3-Flash with oQe quantization. Qwen3.8-Flash-Next combines Lightning MTP with SSD offloading, boosting 4K decode from 22.2 to 58.1 tok/s on M3 Ultra with 96.8% acceptance. The experimental GPU+ANE+CPU path also sees improved stability and long-prompt quality.
More from Infra
- Pure MLX Engine Hits 65 tok/s, Cuts Model Size in Half — EyalToledano · 2026-08-28
- Qwen3.8 Variants: 8-bit Model Cuts Memory by 27% — EyalToledano · 2026-08-28
- 180B-Class Qwen3.8 Model Runs on Just 39GB Memory — EyalToledano · 2026-08-28
- Nvidia's Moat Lies in Scale and Supply Chain Lock-in — firstadopter · 2026-08-28
- Opinion: Qwen's n-gram innovation pops the AI bubble — Acrobatic_Stress1388 · 2026-08-28
- Microsoft Foundry launches AI Gateway for unified traffic and cost governance — davemccollough · 2026-08-28