Struggling to Run DeepSeek Locally on Dual RTX 6000 Ada with vLLM/SGLang

EggDroppedSoup · reddit · 2026-08-03

A developer reports difficulties deploying DeepSeek locally on dual RTX 6000 Ada GPUs. Running the model with dSpark enabled fails on both SGLang and vLLM. While a community NVFP4 quantization had issues, MXFP4 worked better but only with dSpark disabled. The developer is actively seeking configuration recipes.

Original post →

More from Infra

Infra channel →