MLX-Serve v26.9.6: M5 Ultra optimizations push Qwen3.8 decode to 300+ tok/s locally

TheMoonMidas · x · 2026-09-26

ddalcu released mlx-serve v26.9.6 (renamed from MLX Core), with major local inference optimizations targeting M5 Ultra.

Key updates:

Peak on M5 Ultra with Flash Next mixed 4-8 bit and --mtp: 219 tok/s decode, 227 tok/s across 4 streams, 3,150 tok/s prefill on an 8k prompt.

Original post →

More from Infra

Infra channel →