First M5 Ultra benchmarks show 50 tok/s running Qwen3 27B q4 locally
Ashefromapex · reddit · 2026-09-17
A Reddit user spotted the first (unofficial) M5 Ultra benchmarks on omlx.ai: running the q4 quantization of Qwen3 27B at 8k context without MTP, it hits roughly 50 tok/s generation and 1800 tok/s prompt processing. The source's officialness is unclear but the numbers seem reasonable and promising.
More from Infra
- Keeping vLLM's prefix cache warm between agent turns — bolts98 · 2026-09-17
- Running Qwen3.8-Flash-Next on 2x5080: should I move from llama.cpp to vLLM? — whatyathinkk · 2026-09-17
- Nvidia's Jensen Huang says chip sales will double next year — zephyr_z9 · 2026-09-17
- GlobalFoundries and Marvell expand Vermont SiGe capacity for AI optical networking — zephyr_z9 · 2026-09-17
- $26,100 desktop AI datacenter: dual RTX PRO 6000 Blackwell workstation goes open source — dee_hw · 2026-09-17
- Huawei Rumored Scale-Up Node with 4,096 Accelerators Could Pack 384TB of HBM — zephyr_z9 · 2026-09-17