Uzu Inference Engine Launches Speculative Decoding for Qwen3.6 27B

AI lab Uzu released a speculative decoding implementation, its biggest launch to date, with first support for Qwen3.6-27B. The model achieves 105 output tokens per second on an Apple M5 Max with 128GB unified memory.

2026-09-04 ~ 2026-09-04 · 2 related posts