Mirai's Uzu engine hits 105 tok/s with Qwen 27B on M5 Max, 2.9x faster than MLX speculative decoding
awnihannun · x · 2026-09-04
Mirai Labs released speculative decoding in its local inference engine Uzu, starting with Qwen3.6-27B: 105 output tokens/sec entirely on-device on an Apple M5 Max with 128GB unified memory — 2.9x faster than the fastest MLX speculative-decoding implementation they benchmarked. The result comes from full-stack co-design: DFlash plus their Weaver model, tree-based speculative decoding, Mirai quantization, a custom verification algorithm, and Metal kernels optimized for Apple silicon. Their site publishes public benchmarks of uzu vs MLX vs llama.cpp on the same device, using a fixed 1,355-token prompt, with speculative decoding speed macro-averaged over MT-Bench, MATH-500 and HumanEval. The team says tokens/sec is not the metric they ultimately care about, with more to come.
Related event: Uzu Inference Engine Launches Speculative Decoding for Qwen3.6 27B(2 posts)→
More from Infra
- Redditor Compiles Mega List of Open-Source LLM Inference Optimization Projects and Papers — Dramatic-Chard-5105 · 2026-09-04
- Perplexity Portable Computer local runtime now available on Linux for RTX GPUs — cameronstow · 2026-09-04
- Nvidia Launches Personal AI Router for Multi-Server Local Inference Setups — DustNearby2848 · 2026-09-04
- Tensordyne publishes inference whitepaper: TSMC 3nm chip taped out with Broadcom — kimmonismus · 2026-09-04
- Irish data centres now consume 23% of national electricity as residents report mysterious hum — Graham_dePenros · 2026-09-04
- AI buildout bottleneck: US-made wafers still fly to Asia for packaging as capacity grows overseas — ashwinl · 2026-09-04