Run DeepSeek V4-Flash with 1M Context Locally on Dual RTX PRO 6000

dee_hw · x · 2026-08-03

A developer shares an update on building a local AI workstation using dual RTX PRO 6000 GPUs. Thanks to DeepSeek V4-Flash's use of sliding-window attention, the KV cache only holds the recent window. This allows the 192GB dual-GPU rig to run the full 1M context window quietly on a desk, avoiding the need for a loud 4x or 8x server setup.

Related event: DeepSeek-V4-Flash Benchmarked Across Hardware: From Consumer GPUs to DGX Spark(26 posts)→

Original post →

More from Infra

Infra channel →