MTPLX V2 Released: MLX Inference Hits 82 TPS on Mac

YoussofAl · reddit · 2026-07-09

MTPLX V2 has been officially released, introducing a Turbo mode. By utilizing a custom validation-specific quantized matrix multiplication kernel and a compilation validation step, it achieves an 82 TPS inference speed at temperature 0.6 when running the Qwen 3.6 27b model on a MacBook Pro M5 Max.

Additionally, the new version brings major improvements to SSD KV caching and long-context tool calls, alongside the release of preliminary benchmark comparisons.

Related event: MTPLX V2 Boosts Local Mac LLM Inference to 80+ TPS(2 posts)→

Original post →

More from Infra

Infra channel →