MTPLX framework achieves ~2x speed boost for Qwen on Apple Silicon

koc_Z3 · reddit · 2026-08-29

The MTPLX framework claims to significantly accelerate Qwen models on Apple Silicon. Tests on an M1 Max 64GB show Qwen2.5-27B (Q4) decoding at 21 TPS and prefill at up to 111 TPS, a 2x boost. The framework features auto-tuning for draft depth and converts base models to MLX-ready MTP models.

Original post →

More from Infra

Infra channel →