oMLX 0.6.3rc1 Released: Faster ANE Tuner, DFlash 2 Support
alexcovo_eth · x · 2026-08-20
oMLX 0.6.3rc1 is now available with key improvements:
- ANE Tuner V2: Predicts the optimal ANE/GPU split and verifies with at most 3 full runs (down from 11), significantly speeding up tuning.
- Wider Quantization Support: Q5, Q6, and Q8 now work with the hybrid ANE/GPU path. On M3 Ultra with Qwen3.8-27B, Q6 prefill increased from 446 to 561 tok/s.
- DFlash 2 Support: End-to-end runtime support for DFlash 2. Using the z-lab Qwen3.8-27B-DFlash2 draft on M3 Ultra, decode speeds improved by 1.3-1.4x from 4K to 32K context.
- Other fixes include improved distributed-cluster reliability and restored prefix-cache reuse in long sessions.
More from Infra
- Clarification: OpenAI's 20% compute claim refers to monitoring overhead, not total capacity — sjgadler · 2026-08-20
- SkyPilot, VAST Data, and Partners Host AI Infra Meetup — skypilot_org · 2026-08-20
- PolymathicAI Releases The Well: A 15TB Collection of Physics Simulations — tom_doerr · 2026-08-20
- Filesystems beat MCP calls by 100x for context retrieval — ml_guy1 · 2026-08-20
- Addy Osmani on Engineering Roles and AI Agents — addyosmani · 2026-08-20
- AI data center controversy silly? Footprint less than an almond field — justin_hart · 2026-08-20