MLX-Serve 26.9.3 ships: Qwen Flash Next tops 100 tok/s on M4/M5 Max Macs
TheMoonMidas · x · 2026-09-17
MLX-Serve 26.9.3 is out, claiming to be the fastest engine for running local models on Macs — LLMs, video, image, voice cloning, and music generation. Qwen Flash Next now exceeds 100 tok/s on M4/M5 Max with strong prefill speeds. The release is labeled "beta," with the next one focused on polish and refactoring.
More from Infra
- New dashboard sets first public baseline for advanced RL training costs — teortaxesTex · 2026-09-17
- PyTorch Conference NA lineup spotlights torch.compile and custom kernel breakthroughs — PyTorch · 2026-09-17
- Structured-decision trick speeds up DiffusionGemma inference 3-10x with one forward per request — bodonoghue85 · 2026-09-17
- Interactive guide maps the full AI compute stack from grid power to workloads — dr_alphalyrae · 2026-09-17
- How to run Qwen3.8-Flash-Next with N-gram SSD streaming in llama.cpp? — Ambitious_Fold_2874 · 2026-09-17
- How Bell Labs Missed the Microchip: IEEE Spectrum Revisits a Landmark Tech-History Blunder — ArtificialOther · 2026-09-17