Extreme Quantization: DeepSeek-V4-Flash Crushed to 54GB GGUF
giveen · reddit · 2026-08-05
To test the limits of model quantization, a developer aggressively compressed the original 95GB DeepSeek-V4-Flash bf16 model down to 54GB using mixed quantization techniques (an IQ2XXS variant with w2Q2K-AProjQ8-OutQ8).
Performance
- On consumer hardware, this 54GB quantized GGUF build generates text coherently at approximately 20.5 tokens/s.
- Benchmarks show a prompt eval speed of 32.24 tokens/s and a generation speed of 20.54 tokens/s.
Verdict
The author notes that while it's debatable if such extreme 2-bit quantization is practical for complex reasoning, it successfully shaves off over 40GB of memory footprint while maintaining coherent text generation.
More from Infra
- Together AI's Monthly Token Volume Skyrockets from 30B to 400T — togethercompute · 2026-08-05
- API Key Expiry Leads to Runaway Agent, Costs $300 in Idle Compute — voooooogel · 2026-08-05
- Dev Builds Pixel-Art GPU Cluster Dashboard in 20 Mins Using GLM Agent — Porespellar · 2026-08-05
- DeepSeek-V4-Flash Runs with 256k Context on 4x 4090 GPUs — dangerous_inference · 2026-08-05
- Extropic Founder: Running Probabilistic AI on Deterministic Hardware is a 'Demon Tax' — beffjezos · 2026-08-05
- Samsung Showcases Its 3D NAND and HBM5 Memory Technology — BenBajarin · 2026-08-05