oMLX 0.7.0 lands with up to 50% faster long-context decoding and a rebuilt memory guard
pcuenq · x · 2026-10-01
oMLX 0.7.0 brings major speedups on Apple Silicon. Versus the 0.7.0rc1 (oQ4e, Lightning MTP on):
- Qwen3.8-Flash-Next prefill on M5 Max at 16K context: 1,953 → 2,844 tok/s (+46%)
- Qwen3.8-27B decode on M3 Ultra at 64K context: 46.9 → 70.3 tok/s (+50%)
- GLM-5.3-Flash prefill on M3 Ultra at 16K context: 413 → 516 tok/s (+25%)
Other highlights: a completely rebuilt memory guard that safely uses as much RAM as possible, DFlash2 support for GLM-5.3-Flash, faster image follow-ups for Qwen3.8-Flash-Next via a vision feature cache, and fixes for regressions reported against the RC (including ANE prefill and M5 A8 issues).
More from Infra
- Huawei Mate 90 debuts Kirin 9050 Pro, first chip with LogicFolding architecture — ingliguori · 2026-10-01
- NVIDIA A20 Standard Reportedly Skips WMCM Packaging — 'Not Enough Capacity' — zephyr_z9 · 2026-10-01
- TRL hits 1M post-trainings per month, team eyes 1M per week — QGallouedec · 2026-10-01
- SlideDP fine-tunes Qwen2.5-72B on four RTX 4090s, beating FSDP2 throughput by 11.2% — Ruijia Yang · 2026-10-01
- China sets all-time monthly electricity record, but industrial power demand growth slows — teortaxesTex · 2026-10-01
- NVIDIA Dynamo Snapshot cuts LLM serving cold-start times by ~10x — Nasereliver · 2026-10-01