antirez hits 25 t/s running DeepSeek v4.1 Flash across two MacBooks via RDMA
antirez · x · 2026-09-11
antirez continues experimenting with DeepSeek v4.1 Flash on MacBooks:
- Tensor parallelism at 50% per machine over RDMA yields 25 tokens/s, vs 15 t/s with SSD streaming on a single machine.
- He notes the model's decoder is unused during prefill, so on a 128GB machine all layers for large prefills can be fully resident in memory.
- He sees significant optimization headroom for this use case.
More from Infra
- 2.78T-param Kimi K3 runs inference on a single CPU in 8.24 GB of RAM — udmrzn · 2026-09-11
- Keep the Claude Desktop Workflow, Swap in Local Models via Ollama for Privacy — Technovangelist · 2026-09-11
- LithosAI ships Day-0 API inference for DeepSeek-V4.1-Flash at 250+ tokens/s per user — JiaZhihao · 2026-09-11
- TwelveLabs Marengo 3.0 Goes GA in Amazon Bedrock for Video Semantic Search — AWS ML Blog · 2026-09-11
- 1:26 continuous aerial AI video made entirely on a Mac with MiniMax H3 — cocktailpeanut · 2026-09-11
- KV cache gets QAT too: why this model beats others at fp4 KV cache — stochasticchasm · 2026-09-11