antirez hits 25 tok/s running DeepSeek v4.1 Flash across two M5 Max MacBooks via RDMA
antirez · x · 2026-09-11
- antirez reports running DeepSeek v4.1 Flash on a 128GB M5 Max with DwarfStar's SSD streaming: 15 tok/s on a single machine, helped by recent SSD-streaming changes that retain the right experts (and possibly higher expert reuse in DS4.1).
- With tensor parallelism split 50/50 across two MacBooks over RDMA, throughput rises to 25 tok/s. He notes this use case could likely be optimized much further and will push the setup online when ready.
More from Infra
- 2.78T-param Kimi K3 runs inference on a single CPU in 8.24 GB of RAM — udmrzn · 2026-09-11
- Keep the Claude Desktop Workflow, Swap in Local Models via Ollama for Privacy — Technovangelist · 2026-09-11
- LithosAI ships Day-0 API inference for DeepSeek-V4.1-Flash at 250+ tokens/s per user — JiaZhihao · 2026-09-11
- TwelveLabs Marengo 3.0 Goes GA in Amazon Bedrock for Video Semantic Search — AWS ML Blog · 2026-09-11
- 1:26 continuous aerial AI video made entirely on a Mac with MiniMax H3 — cocktailpeanut · 2026-09-11
- KV cache gets QAT too: why this model beats others at fp4 KV cache — stochasticchasm · 2026-09-11