Local Benchmarking of DeepSeek V4 Flash Quantizations: Q3 vs Q8
Spicy_mch4ggis · reddit · 2026-08-04
A developer benchmarked the Unsloth Q8 and Q3 xxs versions of DeepSeek V4 Flash on a 128GB VRAM setup.
The findings reveal that the Q3 version processes prompts 3.5x faster and decodes 2x faster compared to the offloaded Q8 version. The author is evaluating if the speed gains justify the quality loss, and also discusses technical details like speculative decoding, 200k context performance, and thinking budget configurations.
More from Infra
- chapter-tgz adds chapter boundaries to tar.gz for O(1) skipping and parallel reads — charliermarsh · 2026-08-04
- Community-built MiniMax H3 weights target 12GB–24GB GPUs with INT4 and NVFP4 variants — VoidAsuka · 2026-08-04
- Used R940 and two RTX 3090s run DeepSeek V4-Flash at 33 tok/s — AbbreviationsSad5582 · 2026-08-04
- NVIDIA says long-context serving speed is set by architecture choices before training — NVIDIAAI · 2026-08-04
- MiniMax H3 full bf16 run hits 95 GB VRAM and finishes a 15s clip in 2 hours — Moarkush · 2026-08-04
- Windows reset restores RTX 3070 throughput for local Qwen3.6-35B inference — campaigner_ · 2026-08-04