DeepSeek-V4-Flash on RTX 3090 with 128GB RAM: 12.5 tok/s via --n-cpu-moe

Ok_Ninja7526 · reddit · 2026-08-02

Reddit user OkNinja7526 shares how they ran DeepSeek-V4-Flash-0731 UD-IQ3S on an RTX 3090 24GB with 128GB DDR5 overclocked to 5600MHz. By replacing llama.cpp binaries in text-generation-webui and setting the key parameter --n-cpu-moe 39 (keeping some MoE experts in system RAM), they achieved about 12.5 tok/s. The model requires 136GB memory, relying heavily on CPU and RAM bandwidth.

Original post →

More from Infra

Infra channel →