DIY Hybrid Setup Runs DeepSeek Locally, Cutting TTFT from 75s to ~9s

A maker built a 'Frankenstein' hybrid DeepSeek deployment on Threadripper Pro with 192GB VRAM, placing 226 experts on GPU, 165 on CPU, and embeddings on SSD. Sustained optimization cut time-to-first-token from 75 seconds to around 8.9 seconds for 1M-context local inference.

2026-10-10 ~ 2026-10-10 · 3 related posts