DeepSeek Local Deployment: Troubleshooting Severe Speed Drop with Speculative Decoding
Easy_Werewolf7903 · reddit · 2026-08-09
A developer encountered severe performance degradation while locally deploying the DeepSeek-V4-Flash model using llama-server. On a rig with an RTX 4090 and RTX 6000 Pro (120GB VRAM total), using the MTP draft model for speculative decoding yielded a solid 30-40 tokens per second.
However, switching to the DSpark draft model caused generation speeds to plummet to 1-2 t/s. The user ruled out VRAM limitations and provided both complete launch configurations, asking the community for help in identifying the bottleneck in the DSpark parameters.
More from Infra
- Together AI Benchmarks Kimi K3: Ranks #1 Across Multiple Inference Tests — togethercompute · 2026-08-09
- Satire: SD Cards at $2/GB in 2026 Highlights AI Storage Costs — felpix_ · 2026-08-09
- Home Assistant 2026.08 Adds Official llama.cpp Integration — ngxson · 2026-08-09
- Fixing Black Video Outputs with MiniMax H3 on AMD GPUs — Present-Guitar-3967 · 2026-08-09
- Enabling PCIe P2P on Consumer Nvidia GPUs Boosts LLM Throughput by 25% — BidonPomoev · 2026-08-09
- Meituan's LongCat 2.0: Fully Trained and Inferenced on Chinese ASICs — bycloud · 2026-08-09