Struggling to Run DeepSeek Locally on Dual RTX 6000 Ada with vLLM/SGLang
EggDroppedSoup · reddit · 2026-08-03
A developer reports difficulties deploying DeepSeek locally on dual RTX 6000 Ada GPUs. Running the model with dSpark enabled fails on both SGLang and vLLM. While a community NVFP4 quantization had issues, MXFP4 worked better but only with dSpark disabled. The developer is actively seeking configuration recipes.
More from Infra
- From 3D Gaming to AI Dominance: How NVIDIA Seized the Future of Computing — TinfoilTricorn · 2026-08-03
- Defending AI's Thirst: Are Data Centers Really Draining More Water Than Agriculture? — joshwhiton · 2026-08-03
- AI Chip Startup OLIX Raises $312M Series B at $3.3B Valuation — matthewclifford · 2026-08-03
- Run Local LLMs on Mac Easily: llama-macos Offers One-Click Server and WebUI — mervenoyann · 2026-08-03
- DeepSeek V4 Flash Crashes During Prompt Processing on Dual Strix Halo RDMA Setup — WallabyFirm1159 · 2026-08-03
- Raylight Adds Sequence Parallel for MiniMax H3, Halving Generation Time on RTX 2000 ADA — Altruistic_Heat_9531 · 2026-08-03