341GB DeepSeek on a 128GB Mac at 2x speed: shrink the codebase, let an agent optimize
Chida82 · reddit · 2026-10-08
Running 341GB DeepSeek V4.1 Flash (Q2) on a 128GB M5 Max Mac via SSD expert streaming gave only 12 tok/s. The author's breakthrough wasn't a kernel—it was making the codebase small enough for a coding agent to work in.
- Codebase pruning: ds4.c was 85k lines spanning three model families and three GPU backends; he physically deleted everything unneeded (Metal + DeepSeek only), cutting it to 34k lines and the tree from 278k to 150k lines. Optimization experiments went from an afternoon to an hour each.
- Strict verification: every change had to produce bit-identical tokens vs upstream (greedy, 10 prompts) and pass A/B/B/A benchmarks with logits compared bit-for-bit—no KV quant, no approximate kernels.
- Optimizations: async GPU commits across decode layers, thread-pooled expert reads with Metal residency sets, a few kernel fusions per token, prefetching next layer's experts during prefill.
- Dual SSD: a byte-identical GGUF copy on an external Thunderbolt 5 SSD, with prefill reading each layer partly from both drives; startup verification (7s) refuses to run on mismatch.
- Results: decode 12.1→24.4 tok/s (2x), 16k→32k prefill 404→636 tok/s, first token after prefill 2.1–3.0s→0.3–1.1s, zero hardware spend.
More from coding & agent
- a16z backs Preference Model, which open-sources Karotte RL environment framework battle-tested by 1M+ evals — a16z · 2026-10-08
- Every's agent skims meeting notes and only pings you when your name comes up — here's the 4-step setup — every · 2026-10-08
- Exa's setup page swaps dev docs for a copy-paste prompt your coding agent runs — josh_bickett · 2026-10-08
- Haiku 5.5 targets high-volume tasks, works as a coding subagent with Opus/Sonnet — claudeai · 2026-10-08
- Non-coder runs his entire business on an army of Claude Opus 5.5 agents — EXM7777 · 2026-10-08
- Obsidian Starter Kit ships a wiki curator agent that builds sourced knowledge bases — dSebastien · 2026-10-08