DeepSeek V4 Flash 3-bit Quantization Tested on 3x RTX 3090: 119GB VRAM
consultkitapp · reddit · 2026-08-02
A user shared llama-bench benchmark results for the DeepSeek-V4-Flash-0731-UD-Q3KXL quantized model running on 3x RTX 3090 GPUs (Total 72GB VRAM).
Benchmark Data:
- Model Size: 119.40 GiB (approx. 284.33B params).
- Prompt Processing (pp512): 116.04 t/s.
- Text Generation (tg128): 7.71 t/s.
The test utilized 21 GPU offload layers (-ngl 21) and layer split mode (--split-mode layer). The poster noted that the quantization performance could likely be pushed further.
More from Infra
- DeepSeek's New Release Significantly Boosts the Value of Nvidia DGX Spark — firstadopter · 2026-08-02
- Developer Showcases Running Hermes Model Locally on Dell Mini PC — burhop · 2026-08-02
- State of AI Compute Index: Anthropic and OpenAI Shift Heavily to Non-Nvidia Chips — nathanbenaich · 2026-08-02
- Maryland County Passes 18-Month Moratorium on Data Center Construction — LadyGagas913 · 2026-08-02
- Power Shortage Becomes the New Bottleneck for the AI Race Beyond Chips — ingliguori · 2026-08-02
- AMD MI355X vLLM Beats Nvidia B200 on Kimi K2.5 Inference — marksaroufim · 2026-08-02