MTPLX V2 Released: MLX Inference Hits 82 TPS on Mac
YoussofAl · reddit · 2026-07-09
MTPLX V2 has been officially released, introducing a Turbo mode. By utilizing a custom validation-specific quantized matrix multiplication kernel and a compilation validation step, it achieves an 82 TPS inference speed at temperature 0.6 when running the Qwen 3.6 27b model on a MacBook Pro M5 Max.
Additionally, the new version brings major improvements to SSD KV caching and long-context tool calls, alongside the release of preliminary benchmark comparisons.
Related event: MTPLX V2 Boosts Local Mac LLM Inference to 80+ TPS(2 posts)→
More from Infra
- SkyPilot exits stealth with $20M seed round and an AI compute platform for fragmented clouds — jfiance · 2026-07-22
- Nothing phone mockup turns a film joke into a modular design meme — ZeYanjie · 2026-07-22
- Actual Computer says its inference stack is tuned for Nvidia’s consumer Blackwell lineup — markjeffrey · 2026-07-22
- Ben Bajarin says CPU demand is still being badly underestimated — BenBajarin · 2026-07-22
- An energy model says the U.S. could run short of natural gas starting in 2028 — churchkey · 2026-07-22
- Devin adds e2b sandboxes for remote agent execution — badphilosopher · 2026-07-22