Local Benchmarking of DeepSeek V4 Flash Quantizations: Q3 vs Q8

Spicy_mch4ggis · reddit · 2026-08-04

A developer benchmarked the Unsloth Q8 and Q3 xxs versions of DeepSeek V4 Flash on a 128GB VRAM setup.

The findings reveal that the Q3 version processes prompts 3.5x faster and decodes 2x faster compared to the offloaded Q8 version. The author is evaluating if the speed gains justify the quality loss, and also discusses technical details like speculative decoding, 200k context performance, and thinking budget configurations.

Original post →

More from Infra

Infra channel →