Mirai's Uzu engine hits 105 tok/s with Qwen 27B on M5 Max, 2.9x faster than MLX speculative decoding

awnihannun · x · 2026-09-04

Mirai Labs released speculative decoding in its local inference engine Uzu, starting with Qwen3.6-27B: 105 output tokens/sec entirely on-device on an Apple M5 Max with 128GB unified memory — 2.9x faster than the fastest MLX speculative-decoding implementation they benchmarked. The result comes from full-stack co-design: DFlash plus their Weaver model, tree-based speculative decoding, Mirai quantization, a custom verification algorithm, and Metal kernels optimized for Apple silicon. Their site publishes public benchmarks of uzu vs MLX vs llama.cpp on the same device, using a fixed 1,355-token prompt, with speculative decoding speed macro-averaged over MT-Bench, MATH-500 and HumanEval. The team says tokens/sec is not the metric they ultimately care about, with more to come.

Related event: Uzu Inference Engine Launches Speculative Decoding for Qwen3.6 27B(2 posts)→

Original post →

More from Infra

Infra channel →