MLX Fast speeds up Qwen 3.8 27B by 153% on Apple Silicon
alexcovo_eth · x · 2026-08-17
Developer shares progress on MLX Fast optimizing the Qwen 3.8 27B model: under 16 hours, performance improved by 153% over baseline and 2.5x faster than out-of-the-box MTP decode. The tool targets Apple Silicon architecture, challenging the belief that Macs are slow at dense models via speculative decoding.
Related event: MLX Optimization Challenge Speeds Up Qwen 27B on Apple Silicon by Over 150%(4 posts)→
More from Infra
- Stripe to Buy OpenRouter for $7B at 140x Revenue Multiple — aakashgupta · 2026-08-17
- Qwen3.8-27B-Ridge-3.7bpw released, shrinking model size to 11.7GB — udmrzn · 2026-08-17
- Limiting GPU max frequency cuts power from 47W to 23W on DGX Spark cluster with minimal decode impact — MaziyarPanahi · 2026-08-17
- Comfy LTX/H3 VRAM Spike Causes Freezes — Dapper_Astronaut_603 · 2026-08-17
- Piper Sandler maps the 8-layer AI infrastructure stack and key stocks — brucemacv · 2026-08-17
- Deconstructing the financing structures behind massive AI compute deals — demian_ai · 2026-08-17