MLX-Serve 26.9.3 ships: Qwen Flash Next tops 100 tok/s on M4/M5 Max Macs

TheMoonMidas · x · 2026-09-17

MLX-Serve 26.9.3 is out, claiming to be the fastest engine for running local models on Macs — LLMs, video, image, voice cloning, and music generation. Qwen Flash Next now exceeds 100 tok/s on M4/M5 Max with strong prefill speeds. The release is labeled "beta," with the next one focused on polish and refactoring.

Original post →

More from Infra

Infra channel →