DeepSeek-V4-Flash on 4×AMD V620: 300K Context, 21 tok/s Generation

Thin_Pollution8843 · reddit · 2026-08-15

A user runs DeepSeek-V4-Flash-0731 IQ3XXS on 4 AMD Radeon Pro V620 GPUs (128GB VRAM) with DSpark speculative decoding. 32K prompt ingestion at 276 tok/s, short prompt at 379 tok/s, continuous generation at 21 tok/s, accepted generation at 30.6 tok/s. However, q3xxs quantization degrades quality, and VRAM is insufficient for the drafter.

Original post →

More from Infra

Infra channel →