MLX-Serve 27B gets speculative branching, up to 32% faster generation on M4 Max

TheMoonMidas · x · 2026-09-27

Developer ddalcu shares that MLX-Serve now incorporates tricks from TensorFold: it makes several branching guesses for the next tokens and verifies them all in a single pass. On an M4 Max running a 27B model, this yields up to 20% faster generation than TensorFold and 32% versus version 26.9.5 — a notable local-inference throughput gain for the Apple MLX ecosystem.

Original post →

More from Infra

Infra channel →