Running 1M Context on 192GB VRAM: Hybrid DeepSeek Setup Cuts TTFT from 75s to 8s

HankYeomans · x · 2026-10-10

Hank Yeomans built a Frankenstein hybrid DeepSeek inference rig on a Threadripper Pro with 192GB VRAM, fast Kingston Fury CL32 RAM and Samsung 9100 Pro NVMe: 226 experts on GPU, 165 on CPU, and engrams on SSD. After tuning, TTFT for 16K prefill dropped from 75s to 8.9s (7.2s at 32K prefill, 58 tok/s decode), prefill throughput rose from 215 to 1910 tokens/s, and he successfully tested 262K×4 and 1M context windows.

Related event: DIY Hybrid Setup Runs DeepSeek Locally, Cutting TTFT from 75s to ~9s(3 posts)→

Original post →

More from Infra

Infra channel →