Qwen 27B runs 170% faster on Apple Silicon, challenging dense model bottlenecks

gajesh · x · 2026-08-16

Optimizations for running dense models on Apple Silicon have yielded significant results, challenging the belief that Macs are unsuitable for such workloads. Running the Qwen 3.8 27B model, the project achieved a 153% speedup over baseline and 2.5x faster performance compared to out-of-the-box MTP decoding within 16 hours. The improvements leverage speculative decoding and highlight substantial optimization headroom for on-device AI inference.

Related event: Qwen 27B Speed Up Over 150% on Apple Silicon(3 posts)→

Original post →

More from Infra

Infra channel →