27B model tops 100 tok/s on a MacBook via M5-tuned quantization and speculative decoding

teortaxesTex · x · 2026-09-09

The norpadon team co-designed quantization and speculative decoding around Apple's M5 neural accelerators, getting a 27B model to run at over 100 tokens/s on a MacBook — for now on M5+ chips only. The latest version also shows significantly improved performance, especially on prefill.

Original post →

More from Infra

Infra channel →