Hobbyist runs hybrid GPU/CPU/SSD MoE DeepSeek setup, cutting TTFT from 75s to 8.9s

HankYeomans · x · 2026-10-10

The author built a "Frankenstein Deepseek v4.1" heterogeneous local deployment: 226 experts on GPU, 165 experts on CPU, and engrams on SSD. Through iterative optimization, TTFT dropped from 75s to 8.9s, prefill throughput rose from 215 to 1910 tokens/sec, and context window was extended from 32k to successfully tested 262k×4 and 1M configurations. A great learning project, they say.

Related event: DIY Hybrid Setup Runs DeepSeek Locally, Cutting TTFT from 75s to ~9s(3 posts)→

Original post →

More from Infra

Infra channel →