DeepSeek Hits 40 tok/s Locally on M3 Ultra Mac Studio
zephyr_z9 · x · 2026-08-01
A developer ran the latest DeepSeek model on an M3 Ultra Mac Studio with 512GB of memory, achieving an inference speed of 40 tokens/second. Commenters noted that this provides a locally runnable experience approaching the capabilities of Claude 3 Opus 4.7-4.8, making it highly usable.
More from Infra
- SGLang Supports Inkling-Small on Dual DGX Spark, Hits 24 tok/s — ying11231 · 2026-08-01
- macmon: Open-Source Terminal Performance Monitor for Apple Silicon Hits 1.8k Stars — tom_doerr · 2026-08-01
- Scaling Kimi K3 on H200s: Engineering Insights from 1000+ Chips — hsu_byron · 2026-08-01
- Laguna Doubles Performance: Significant Mac Inference Speedup Without Speculative Decoding — gajesh · 2026-08-01
- Will the AI Agent Explosion Overload and Break Internet Infrastructure? — Ok-Video4323 · 2026-08-01
- Kimi K3 Hits Record 172 Tokens/sec in Inference Speed — AccBalanced · 2026-08-01