DeepSeek Local Deployment: Troubleshooting Severe Speed Drop with Speculative Decoding

Easy_Werewolf7903 · reddit · 2026-08-09

A developer encountered severe performance degradation while locally deploying the DeepSeek-V4-Flash model using llama-server. On a rig with an RTX 4090 and RTX 6000 Pro (120GB VRAM total), using the MTP draft model for speculative decoding yielded a solid 30-40 tokens per second.

However, switching to the DSpark draft model caused generation speeds to plummet to 1-2 t/s. The user ruled out VRAM limitations and provided both complete launch configurations, asking the community for help in identifying the bottleneck in the DSpark parameters.

Original post →

More from Infra

Infra channel →