Uzu Inference Engine Launches Speculative Decoding for Qwen3.6 27B
AI lab Uzu released a speculative decoding implementation, its biggest launch to date, with first support for Qwen3.6-27B. The model achieves 105 output tokens per second on an Apple M5 Max with 128GB unified memory.
2026-09-04 ~ 2026-09-04 · 2 related posts
- Mirai's Uzu engine hits 105 tok/s with Qwen 27B on M5 Max, 2.9x faster than MLX speculative decoding — awnihannun · 2026-09-04
- Uzu lab ships speculative decoding implementation, launching with Qwen3.6 27B support — awnihannun · 2026-09-04