DeepSeek Hits 40 tok/s Locally on M3 Ultra Mac Studio

zephyr_z9 · x · 2026-08-01

A developer ran the latest DeepSeek model on an M3 Ultra Mac Studio with 512GB of memory, achieving an inference speed of 40 tokens/second. Commenters noted that this provides a locally runnable experience approaching the capabilities of Claude 3 Opus 4.7-4.8, making it highly usable.

Original post →

More from Infra

Infra channel →