Hobbyist runs hybrid GPU/CPU/SSD MoE DeepSeek setup, cutting TTFT from 75s to 8.9s
HankYeomans · x · 2026-10-10
The author built a "Frankenstein Deepseek v4.1" heterogeneous local deployment: 226 experts on GPU, 165 experts on CPU, and engrams on SSD. Through iterative optimization, TTFT dropped from 75s to 8.9s, prefill throughput rose from 215 to 1910 tokens/sec, and context window was extended from 32k to successfully tested 262k×4 and 1M configurations. A great learning project, they say.
Related event: DIY Hybrid Setup Runs DeepSeek Locally, Cutting TTFT from 75s to ~9s(3 posts)→
More from Infra
- Starship Could Cut Cost to Orbit 100x to ~$185K Per Ton, Unlocking New Businesses — claud_fuen · 2026-10-10
- DuckDB v2.0 CLI agent mode cuts agent-read tokens by 59% on TPC-H benchmarks — josh_wills · 2026-10-10
- Datology releases Zephon, a deterministic on-the-fly dataloader born from MosaicML Streaming's legacy — josh_wills · 2026-10-10
- Tsinghua's TokenRouter: Token-Level LLM Routing Hits Up to 64.15X Serving Throughput — rohanpaul_ai · 2026-10-10
- Meta Muse Auto-Routes to OpenRouter Free Models for Zero-Cost Long Tasks — sven_ai · 2026-10-10
- Nvidia CEO's son-in-law becomes VP as Apple cuts iPhone 18 Pro orders 15%+ — 创业邦 · 2026-10-10