Running DeepSeek-V4-Flash Extreme Quantization on Single RTX 3090

nikhilprasanth · reddit · 2026-08-03

A Reddit user shared a detailed test of running the extreme quantized version of DeepSeek-V4-Flash (IQ2XS-Experts-Q80) on a single RTX 3090. The setup uses 24GB VRAM and 128GB RAM via llama.cpp. Despite heavy compression, the model successfully generated structurally complete and usable code, losing only some fine details.

Original post →

More from Infra

Infra channel →