DeepSeek V4 Flash 3-bit Quantization Tested on 3x RTX 3090: 119GB VRAM

consultkitapp · reddit · 2026-08-02

A user shared llama-bench benchmark results for the DeepSeek-V4-Flash-0731-UD-Q3KXL quantized model running on 3x RTX 3090 GPUs (Total 72GB VRAM).

Benchmark Data:

The test utilized 21 GPU offload layers (-ngl 21) and layer split mode (--split-mode layer). The poster noted that the quantization performance could likely be pushed further.

Original post →

More from Infra

Infra channel →