Running local LLMs on an M5 Pro 64GB: ~20 tok/s on Qwen 27B, DS4 Flash at just 8 tok/s
rJohn420 · reddit · 2026-09-13
A Reddit user shares real-world local inference numbers on an M5 Pro with 64GB RAM: around 20 tok/s sustained on Qwen 3 27B (about 30 tok/s in the first 2-4k tokens), but heavy thinking makes it feel slow. Trying antirez's DS4, DeepSeek V4 Flash drops to roughly 8 tok/s sustained — effectively unusable. The thread asks what others run on the same hardware.
More from Infra
- Chinese Optical Transceiver Makers Dodge US Blacklist; Zhongji Innolight Jumps 4% — pstAsiatech · 2026-09-14
- Chinese team boosts ferroelectric memory endurance 100x, exceeding 10B write cycles in Science — pstAsiatech · 2026-09-14
- Running CUDA workloads on AMD GPUs under Windows sparks ROCm discussion — Worried_Ad_1816 · 2026-09-14
- Reddit speculates DeepSeek-style KVCache compression could let 32GB GPUs run 54B-class models — pmttyji · 2026-09-14
- Google says AI server payback period is under 2 years, half that for its own silicon — Beth_Kindig · 2026-09-13
- 13 attention mechanisms AI engineers must know, organized by the bottleneck each solves — blaizedsouza · 2026-09-13