DeepSeek-V4-Flash on RTX 3090 with 128GB RAM: 12.5 tok/s via --n-cpu-moe
Ok_Ninja7526 · reddit · 2026-08-02
Reddit user OkNinja7526 shares how they ran DeepSeek-V4-Flash-0731 UD-IQ3S on an RTX 3090 24GB with 128GB DDR5 overclocked to 5600MHz. By replacing llama.cpp binaries in text-generation-webui and setting the key parameter --n-cpu-moe 39 (keeping some MoE experts in system RAM), they achieved about 12.5 tok/s. The model requires 136GB memory, relying heavily on CPU and RAM bandwidth.
More from Infra
- Kernel-level optimizations for DeepSeek and Laguna models on Apple Silicon — gajesh · 2026-08-02
- Apple Silicon inference runs 137% faster as MLX Challenge pushes edge limits — gajesh · 2026-08-02
- NVIDIA's Rubin GPUs Expected to Cut Inference Costs by 90% by Late 2026 — haider1 · 2026-08-02
- Choosing a GPU Cloud Provider for Production: Beyond Price — 9ds996Dev · 2026-08-02
- Rust Compiler Performance: rustdoc Build Time Slashed by 28% — charliermarsh · 2026-08-02
- Compute Still King: French AI Circle Reflects on Efficiency Gap with DeepSeek — AymericRoucher · 2026-08-02