MTPLX boosts Apple Silicon local inference speed by 3x

julianharris · x · 2026-08-25

MTPLX is a native Mac app and CLI tool for Apple Silicon that leverages the Multi-Token Prediction (MTP) heads found in modern models like Qwen 3.8 to accelerate local inference by approximately 3x. Benchmarks indicate that an M3 running MTPLX with Qwen 3.8 27B outperforms an unoptimized M3 Max, which struggles with slow prefill speeds despite handling larger models.

Original post →

More from Infra

Infra channel →