Running DeepSeek Locally on MacBook Pro Hits Nearly 40 tokens/s
victormustar · x · 2026-08-03
A developer successfully achieved high-efficiency local inference for the DeepSeek model on an M5 Max MacBook Pro. Even with a context length exceeding 15,000 tokens, the generation speed reached 36 to 55 tokens/s.
This impressive performance is largely attributed to combining @unsloth's UD-Q2KXL quantization with @ggmlorg's MXFP4 DSpark draft model for speculative decoding.
More from Infra
- Running MiniMax H3 Locally: Extremely High VRAM and RAM Usage Reported — Full_Astronomer_5438 · 2026-08-03
- Cloudflare Launches Billable Usage API for Programmatic Cost Visibility — ritakozlov · 2026-08-03
- Cloudflare Workers & Containers Add Inbound TCP and gRPC Support — ritakozlov · 2026-08-03
- Cloudflare Launches @cloudflare/computer: A Dedicated Runtime Environment for Every Agent — threepointone · 2026-08-03
- Turso Database Overcomes SQLite Limits with Concurrent Writes — glcst · 2026-08-03
- Cloudflare Details Inference Optimizations for Running Kimi and GLM at Scale — michellechen · 2026-08-03