Qwen 3.8 27B on M4 Max: 72.1 tok/s for code via ANE + MTP

TheMoonMidas · x · 2026-08-22

A developer shares an optimized recipe for running Qwen 3.8 27B on a Mac Studio M4 Max. By using the oMLX 0.6.3rc2 framework combined with ANE (Apple Neural Engine) for prefill acceleration and native Multi-Token Prediction (MTP, k=3), significant performance gains were achieved.

This represents an 11% increase in decode speed over the previous best method. All tests were run locally without cloud APIs, and the raw benchmark data and configuration have been open-sourced on GitHub.

Original post →

More from Infra

Infra channel →