Running DeepSeek V4-Flash Locally: Dual RTX 6000 Rig Handles Only Single User
dee_hw · x · 2026-08-02
A developer shared hands-on experience running DeepSeek V4-Flash on a 2× RTX PRO 6000 build.
- VRAM limits: VRAM caps the context at 131K. This means if two requests each exceed 65K context, one has to wait for the other to finish.
- Context usage: A coding task typically runs >60K context, an agentic task 30–40K, and a chat app 10–15K.
- Verdict: This rig is great for a single personal user and fine for a small team, but it is not suitable as a production box.
Related event: DeepSeek V4-Flash Tested on Dual RTX 6000(2 posts)→
More from Infra
- Building a 2400W Multi-GPU Workstation: Reliability of Dual PSU Setups — Generic_Name_Here · 2026-08-02
- Exploring Local AI Video Generation: What Are the Limits of RAM Offloading? — Independent-Frequent · 2026-08-02
- 4-bit KV Cache Tested: Perplexity Surges 43% in Long Contexts, q8_0 is the Sweet Spot — Dhan295 · 2026-08-02
- Optimizing AI for Edge: AIMET Uses Quantization and Pruning for On-Device Deployment — carrycooldude · 2026-08-02
- Bought a $5k Mac Studio for local LLMs, ended up running hundreds of subagents — EverydayAI_ · 2026-08-02
- DeepSeek's New Release Significantly Boosts the Value of Nvidia DGX Spark — firstadopter · 2026-08-02