456GB DeepSeek v4.1 runs locally at 40 tok/s with Threadripper + dual RTX 6000 hybrid setup
HankYeomans · x · 2026-09-22
A developer got DeepSeek v4.1 MXFP4 (about 456GB) running locally at 40 tok/s using a hybrid setup: 2x RTX 6000 Pro Max-Q for hot experts, a Threadripper Pro with 256GB CL32/6400MT RAM handling spillover experts, and Samsung 9100 NVMe for Engrams.
The goal was pushing non-GPU hardware to its limits and showing what big-memory workstation platforms can do for local LLM inference.
More from Infra
- Is upgrading from 2x to 4x RTX 3090 worth it for local LLM work? — fgoricha · 2026-09-22
- Qwen-Image local on a 24GB MacBook Pro takes 5-6 minutes per image — vista8 · 2026-09-22
- Cerebras CEO on Jensen Huang: a decade trading as 'nobody' before Nvidia made it — rohanpaul_ai · 2026-09-22
- How Tencent Hunyuan packed a 770B model into 214 GiB with 5-bit-per-4-weights quantization — TencentHunyuan · 2026-09-22
- Agent Substrate roadmap: sub-second suspend/resume runtime for dense agent deployments — rakyll · 2026-09-22
- StepFun's Step 5 Preview scores 44 on AA Intelligence Index at ~2.8x lower cost than Kimi K3 — ArtificialAnlys · 2026-09-22