DIY hybrid GPU/CPU/SSD rig cuts DeepSeek TTFT from 75s to 8.9s at 16K prefill
HankYeomans · x · 2026-10-10
The author built a "Frankenstein" tiered inference setup running DeepSeek (self-described v4.1): 226 experts on GPU, 165 on CPU, and engrams on SSD. After optimization, TTFT for 16K prefill dropped from 75s to 8.9s and prefill throughput rose from 215 to 1910 tokens/s; context window grew from 32K to successful tests of 262K×4 up to 1M.
Related event: DIY Hybrid Setup Runs DeepSeek Locally, Cutting TTFT from 75s to ~9s(3 posts)→
More from Infra
- DuckDB v2.0 CLI agent mode cuts agent-read tokens by 59% on TPC-H benchmarks — josh_wills · 2026-10-10
- Datology releases Zephon, a deterministic on-the-fly dataloader born from MosaicML Streaming's legacy — josh_wills · 2026-10-10
- Tsinghua's TokenRouter: Token-Level LLM Routing Hits Up to 64.15X Serving Throughput — rohanpaul_ai · 2026-10-10
- Meta Muse Auto-Routes to OpenRouter Free Models for Zero-Cost Long Tasks — sven_ai · 2026-10-10
- vLLM thread (4/5): locality-domain MoE sharding speeds up decode 1.2x — vllm_project · 2026-10-10
- vLLM lands NVIDIA Vera Rubin support, hitting 7.8x GB200 throughput on MiniMax M3 — vllm_project · 2026-10-10