Running DeepSeek V4 Flash on a Single GPU

live4evrr · reddit · 2026-07-10

A developer successfully ran DeepSeek V4 Flash on a single RTX 6000 Pro using a custom version of vLLM-Moet, claiming a 130K context length can still be set on a single card. The post notes that around 150GB of system RAM is required to load the safetensor shards, but once loaded, the model fits entirely into VRAM for inference.

Original post →

More from Infra

Infra channel →