DeepSeek V4 Flash Local Tested

kevin_1994 · reddit · 2026-07-11

A sharing of local running experiences with DeepSeek V4 Flash on consumer hardware.

The author's setup:

After testing various quantization and parameter combinations, it was successfully run using unsloth's UD-Q2KXL quantization, yielding these speeds:

Key observations:

In comparison, the author finds it slightly "smarter" but slower than Qwen 3.6 27B Q4KXL; however, because it exhibits less "overthinking", the total time for many tasks is actually similar. The author believes that once issues like flash attention, batch/microbatch, and context quantization are fixed, this model will become much more practical on 4090/3090 GPUs.

Original post →

More from Infra

Infra channel →