MTPLX V2 Released: Runs Qwen 27B on MacBook M5 Max at 72+ Tokens/sec

awnihannun · x · 2026-07-08

MTPLX V2 has been released, a local inference tool based on the MLX framework and optimized specifically for Apple Silicon. In real-world testing, running Qwen 3.6 27B on a MacBook Pro M5 Max achieved 72+ tokens/sec. The author claims this is currently the fastest way to run models on MLX.

Related event: MTPLX V2 Boosts Local Mac LLM Inference to 80+ TPS(2 posts)→

Original post →

More from Infra

Infra channel →