Mirai's uzu engine brings speculative decoding to Apple M5, hitting 105 tok/s on Qwen3.6 27B
TheMoonMidas · x · 2026-09-04
Mirai released a speculative decoding implementation in its local inference engine uzu, initially supporting Qwen3.6 27B with Qwen3.8 27B and Muse Glimmer to follow.
- On Apple M5 Max, Qwen3.6 27B 4-bit (Mirai-M) outputs 105 tok/s — nearly 2x MTPLX (55 tok/s) and over 3x llama.cpp (30 tok/s) at comparable quantization.
- Gains vary by domain: 113.84 tok/s on HumanEval and 117.30 tok/s on MATH-500 vs 84 tok/s on MT-Bench, since code is easier to predict than chat.
- The draft model, quantization format, verification algorithm, and GPU kernels are co-designed from the ground up for Apple M5's Neural Accelerators.
- Runs via brew install mirai (14.5 GB checkpoint), with an SDK for Swift, TypeScript, Python, and Rust.
More from Infra
- Nunchux launches: one API aggregating 30+ image, video and avatar models — junyanz89 · 2026-09-04
- Inco AI enters public beta, tops Artificial Analysis output speed for 4 open models — Lianhuiq · 2026-09-04
- Starship fires all 33 Raptors at once, hitting 74M newtons—double Saturn V — PeterDiamandis · 2026-09-04
- xAI wins approval for fifth Memphis-area data center in $40M land swap with Southaven — chrisgrayson · 2026-09-04
- Best local models for 12GB of VRAM: Gemma-4-12B remains the pick — GlennCameronjr · 2026-09-04
- Rumor: When One Major AI Service Goes Down, Its Traffic Takes the Rest Down With It — jxnlco · 2026-09-04