Running local LLMs on an M5 Pro 64GB: ~20 tok/s on Qwen 27B, DS4 Flash at just 8 tok/s

rJohn420 · reddit · 2026-09-13

A Reddit user shares real-world local inference numbers on an M5 Pro with 64GB RAM: around 20 tok/s sustained on Qwen 3 27B (about 30 tok/s in the first 2-4k tokens), but heavy thinking makes it feel slow. Trying antirez's DS4, DeepSeek V4 Flash drops to roughly 8 tok/s sustained — effectively unusable. The thread asks what others run on the same hardware.

Original post →

More from Infra

Infra channel →