First M5 Ultra benchmarks show 50 tok/s running Qwen3 27B q4 locally

Ashefromapex · reddit · 2026-09-17

A Reddit user spotted the first (unofficial) M5 Ultra benchmarks on omlx.ai: running the q4 quantization of Qwen3 27B at 8k context without MTP, it hits roughly 50 tok/s generation and 1800 tok/s prompt processing. The source's officialness is unclear but the numbers seem reasonable and promising.

Original post →

More from Infra

Infra channel →