MLX Fast speeds up Qwen 3.8 27B by 153% on Apple Silicon

alexcovo_eth · x · 2026-08-17

Developer shares progress on MLX Fast optimizing the Qwen 3.8 27B model: under 16 hours, performance improved by 153% over baseline and 2.5x faster than out-of-the-box MTP decode. The tool targets Apple Silicon architecture, challenging the belief that Macs are slow at dense models via speculative decoding.

Related event: MLX Optimization Challenge Speeds Up Qwen 27B on Apple Silicon by Over 150%(4 posts)→

Original post →

More from Infra

Infra channel →