Used R940 and two RTX 3090s run DeepSeek V4-Flash at 33 tok/s
AbbreviationsSad5582 · reddit · 2026-08-04
DeepSeek V4-Flash runs on a used R940 and two RTX 3090s at 33 tok/s
The post benchmarks the official DeepSeek V4-Flash-0731 checkpoint on commodity used hardware, arguing that capacity matters more than raw GPU speed for this sparse MoE model.
- Model size: 284B total / 13B active, with an official checkpoint footprint of 156 GB.
- The setup uses a Dell PowerEdge R940 with 4× Xeon Platinum 8268 CPUs, 768 GB DDR4-2933, and 2× RTX 3090 24 GB GPUs.
- The inference stack is Lvllmds4-x v2.3.8, a vLLM fork with lkmoe v2.3.1 doing NUMA-aware CPU-GPU hybrid MoE execution.
- Because Ampere lacks native FP4/FP8 compute, the fork routes weights through Marlin kernels; the checkpoint is quantization-aware trained and ships routed experts in MXFP4.
- Reported throughput: 33 tok/s single stream, 53–68 tok/s aggregate at 4 concurrent users, and 47–63 tok/s at 8 concurrent users.
- Power draw was about 1,000 W under decode, with the two 3090s only drawing 136–145 W on average; the author estimates roughly $94/month for 24/7 operation at $0.13/kWh.
- The main claim is that a used enterprise server plus second-hand GPUs can run the full checkpoint for around $6K all-in.
More from Infra
- Nuclear Startup Valar Raises $1B Led by Sequoia to Scale Reactors — kleffew94 · 2026-08-04
- AI API revenue still trails hyperscaler capex by a wide margin in 2025 chart — SurpriseDog9000 · 2026-08-04
- U.S. heartland backlash grows as AI data centers reshape local communities — altryne · 2026-08-04
- Semiconductors and data centers are being built far slower than AI demand — robleclerc · 2026-08-04
- Gemma 4 31B can use over 13× more KV-cache memory than DeepSeek V4 Flash — teortaxesTex · 2026-08-04
- MCP server brings structured compile, flash and stateful GDB to embedded boards — Historical_Court795 · 2026-08-04