Running DeepSeek-V4-Flash Extreme Quantization on Single RTX 3090
nikhilprasanth · reddit · 2026-08-03
A Reddit user shared a detailed test of running the extreme quantized version of DeepSeek-V4-Flash (IQ2XS-Experts-Q80) on a single RTX 3090. The setup uses 24GB VRAM and 128GB RAM via llama.cpp. Despite heavy compression, the model successfully generated structurally complete and usable code, losing only some fine details.
More from Infra
- ComfyUI Node Test: Sage Attention Significantly Accelerates MiniMax H3 — ConstructionOdd7870 · 2026-08-03
- Big Tech's GPU Hoarding Raises Open-Weight Hosting Costs Near Closed-Model Levels — xiaosun86 · 2026-08-03
- UK Sovereign AI Invests in OLIX to Rebuild AI Infrastructure From First Principles — HZoete · 2026-08-03
- EPA Rules Power for Data Centers Can Sidestep Pollution Laws — KeanuRave100 · 2026-08-03
- NVIDIA RTX 50 Series GPU Prices Surge Up to 30% in Korea Amid Memory Cost Hikes — basedjensen · 2026-08-03
- Running MiniMax H3 on Laptop: 16GB VRAM Renders 5s Video in 3 Minutes — robomar_ai_art · 2026-08-03