DeepSeek Runs at Only 8 tps on Mac Studio

BitXorBit · reddit · 2026-07-08

Someone locally ran DeepSeek V4 Flash on a Mac Studio M3 Ultra, noting that llama.cpp now supports the model and unsloth has released a GGUF version. The main takeaway is that the author's real-world test yielded only about 8 token/s. They shared the llama-server launch parameters, asking if this speed is normal or if there is a configuration issue.

Original post →

More from Infra

Infra channel →