M5 Max user gets ~20 tok/s running DeepSeek locally, asks which open models to pick
A_Wild_Entei · reddit · 2026-09-06
An M5 Max owner shares local inference numbers: antirez's ds4 runs at about 20 tok/s on his machine, which was fine for his needs. Noting recent progress in Qwen, the DeepSeek V4 vision model, and GLM, he asks which quant levels to use and what the best local model/speed combo is right now. The thread is a practical discussion for running open-source LLMs on Apple Silicon.
More from Infra
- T-Glass shortage worsens: Kinsus losing 10-15% of monthly ABF revenue, 25% capacity expansion planned for 2027 — zephyr_z9 · 2026-09-06
- Bump-less 3D stacking goes practical: Intel Diamond Rapids first, AMD Zen rumored next — bookwormengr · 2026-09-06
- Self-hosting AI: what rigs do local LLM runners actually use? — Fun_Kangaroo512 · 2026-09-06
- Reddit debate: is compute the real hurdle to automating all cognitive labour? — Vivid-Flamingo-644 · 2026-09-06
- 200 tok/s on 8GB VRAM: dev benchmarks 6 small models for local AI — TheMoonMidas · 2026-09-06
- Reddit debate: is compute the real hurdle to automating all cognitive labour? — Vivid-Flamingo-644 · 2026-09-06