DIY hybrid GPU/CPU/SSD rig cuts DeepSeek TTFT from 75s to 8.9s at 16K prefill

HankYeomans · x · 2026-10-10

The author built a "Frankenstein" tiered inference setup running DeepSeek (self-described v4.1): 226 experts on GPU, 165 on CPU, and engrams on SSD. After optimization, TTFT for 16K prefill dropped from 75s to 8.9s and prefill throughput rose from 215 to 1910 tokens/s; context window grew from 32K to successful tests of 262K×4 up to 1M.

Related event: DIY Hybrid Setup Runs DeepSeek Locally, Cutting TTFT from 75s to ~9s(3 posts)→

Original post →

More from Infra

Infra channel →