MLX-Serve v26.9.6: M5 Ultra optimizations push Qwen3.8 decode to 300+ tok/s locally
TheMoonMidas · x · 2026-09-26
ddalcu released mlx-serve v26.9.6 (renamed from MLX Core), with major local inference optimizations targeting M5 Ultra.
Key updates:
- M5 Ultra: Qwen3.8 Flash Next reaches 243-300+ tok/s decode and 3,362 tok/s prefill; fused MoE kernel adds +14% on verification, and fused multi-stream verification adds +13% at 4 streams
- Laya typed decisions: new POST /v1/decisions endpoint with local-JEV support
- MTP drafts from the conversation itself when replies copy it, speeding file rewrites/edits by 1.3-1.5x (on by default; disable with MLXSERVEMTPLOOKUP=0)
- GDN layers now run as one GPU dispatch, helping every Mac a few percent
- Fixed media and preservethinking handling, UI updates
Peak on M5 Ultra with Flash Next mixed 4-8 bit and --mtp: 219 tok/s decode, 227 tok/s across 4 streams, 3,150 tok/s prefill on an 8k prompt.
More from Infra
- SemiAnalysis probes DeepSeek V4.1 Flash's Engram gates, revealing how the model activates text patterns — teortaxesTex · 2026-09-26
- FreeToken: Open-Source Engine Runs 290B+ MoE Models Locally on Gaming PCs, 13.8k Stars — solyarisoftware · 2026-09-26
- AI buildout needs ~$5.5 per GPU-hour revenue, within current pricing: Columbia report — rohanpaul_ai · 2026-09-26
- SemiAnalysis tears down Apple M6 and TSMC N2 for free, revealing GAAFET scaling details — AIFlow_ML · 2026-09-26
- llama.cpp PR claims 3-7x faster CPU prompt processing via VNNI — jacek2023 · 2026-09-26
- AxonDAO launches AxonOS: a GPU-native Linux desktop for science in your browser — melnykowycz · 2026-09-26