Running DeepSeek Locally on MacBook Pro Hits Nearly 40 tokens/s

victormustar · x · 2026-08-03

A developer successfully achieved high-efficiency local inference for the DeepSeek model on an M5 Max MacBook Pro. Even with a context length exceeding 15,000 tokens, the generation speed reached 36 to 55 tokens/s.

This impressive performance is largely attributed to combining @unsloth's UD-Q2KXL quantization with @ggmlorg's MXFP4 DSpark draft model for speculative decoding.

Original post →

More from Infra

Infra channel →