Qwen3.6 27B Achieves 81 TPS Local Inference on M5 Max
max_paperclips · x · 2026-07-03
Shared test data reveals that running the Qwen 3.6 27B model on an Apple M5 Max chip can achieve an inference speed of 81 tokens/second, maintaining 63 TPS even when generating 11k tokens. This test was conducted using the MTPLX inference optimization tool, whose upcoming V2 release will further enhance local LLM inference performance.
More from Infra
- Intel 10-Q points to 18A/14A progress and “potential significant external customers” — BenBajarin · 2026-07-27
- Moonshot’s Kimi K3 lands on Together with reserved throughput and 65% lower cost — togethercompute · 2026-07-27
- OpenAI may be hitting compute limits as Codex and ChatGPT Work jump from 2M to 10M users — JoshuaJBouw · 2026-07-27
- NVIDIA says Vera CPU is speeding up next-gen CPU and GPU design cycles — nordicinst · 2026-07-27
- NVIDIA says Nemotron 3 Ultra hit 97.1% on agentic RTL chip-design tasks — NVIDIAAI · 2026-07-27
- NVIDIA says Vera CPU lifted selected EDA workloads by up to 1.5x — NVIDIA Blog · 2026-07-27