Running 1M Context on 192GB VRAM: Hybrid DeepSeek Setup Cuts TTFT from 75s to 8s
HankYeomans · x · 2026-10-10
Hank Yeomans built a Frankenstein hybrid DeepSeek inference rig on a Threadripper Pro with 192GB VRAM, fast Kingston Fury CL32 RAM and Samsung 9100 Pro NVMe: 226 experts on GPU, 165 on CPU, and engrams on SSD. After tuning, TTFT for 16K prefill dropped from 75s to 8.9s (7.2s at 32K prefill, 58 tok/s decode), prefill throughput rose from 215 to 1910 tokens/s, and he successfully tested 262K×4 and 1M context windows.
Related event: DIY Hybrid Setup Runs DeepSeek Locally, Cutting TTFT from 75s to ~9s(3 posts)→
More from Infra
- Fireworks AI confirms security incident involving unauthorized use of internal credentials — lqiao · 2026-10-10
- ChapterPal Brings Offline Gemini Nano AI Tutor to Android, iOS Version Coming Soon — burkov · 2026-10-10
- Chat with AI could feel outdated in 1-2 years as agent swarms push traffic 1,000x — dumpshoot · 2026-10-10
- Starship Could Cut Cost to Orbit 100x to ~$185K Per Ton, Unlocking New Businesses — claud_fuen · 2026-10-10
- Data center darling's $30B IPO dream crushed in 48 hours — mfiguiere · 2026-10-10
- DuckDB v2.0 CLI agent mode cuts agent-read tokens by 59% on TPC-H benchmarks — josh_wills · 2026-10-10