Run DeepSeek V4-Flash with 1M Context Locally on Dual RTX PRO 6000
dee_hw · x · 2026-08-03
A developer shares an update on building a local AI workstation using dual RTX PRO 6000 GPUs. Thanks to DeepSeek V4-Flash's use of sliding-window attention, the KV cache only holds the recent window. This allows the 192GB dual-GPU rig to run the full 1M context window quietly on a desk, avoiding the need for a loud 4x or 8x server setup.
More from Infra
- NVIDIA's Vera CPU Architecture: Targeting GPU Training and Agent Sandboxes — zephyr_z9 · 2026-08-03
- Citadel Forecasts Over $500B in New Debt to Finance AI Chips by 2028 — Polymarket · 2026-08-03
- IT Engineer Shares 6-Month Review of 256GB VRAM 'Data Center on Wheels' — SweetHomeAbalama0 · 2026-08-03
- Interactive Explainer: Visualizing the Journey of a Data Center Request — thisiskp_ · 2026-08-03
- AI Data Centers Consume 1.5B Gallons of Water Annually, Sparking Environmental Debate — ZeroStateReflex · 2026-08-03
- RTX 5090 benchmark: Minimax H3 takes 16 minutes for an 11-second video — BlackBeardAI · 2026-08-03