Mirai's uzu engine brings speculative decoding to Apple M5, hitting 105 tok/s on Qwen3.6 27B

TheMoonMidas · x · 2026-09-04

Mirai released a speculative decoding implementation in its local inference engine uzu, initially supporting Qwen3.6 27B with Qwen3.8 27B and Muse Glimmer to follow.

Related event: Uzu inference engine adds speculative decoding, 3x faster than llama.cpp on M5(3 posts)→

Original post →

More from Infra

Infra channel →