Qwen3.8-Flash on 4x AMD V620 hits 1,300 PP and 70+ tok/s on coding

Thin_Pollution8843 · reddit · 2026-09-17

A Reddit user benchmarked the quantized Qwen3.8-Flash-Next (W4A16-AutoRound, roughly between q5-xl and q6-xl quality) on 4 RDNA2 AMD V620 GPUs via the community vllm-rdna project: 1,300-1,393 tok/s prompt processing and 56-59 tok/s generation at 32K-128K contexts, rising to 68-72.5 tok/s on coding suites. A budget path to local inference on aging datacenter GPUs.

Original post →

More from Infra

Infra channel →