M5 Max user gets ~20 tok/s running DeepSeek locally, asks which open models to pick

A_Wild_Entei · reddit · 2026-09-06

An M5 Max owner shares local inference numbers: antirez's ds4 runs at about 20 tok/s on his machine, which was fine for his needs. Noting recent progress in Qwen, the DeepSeek V4 vision model, and GLM, he asks which quant levels to use and what the best local model/speed combo is right now. The thread is a practical discussion for running open-source LLMs on Apple Silicon.

Original post →

More from Infra

Infra channel →