Qwen on 4xR9700: 120 t/s Gen and 12k t/s Prefill with Optimized vLLM

sloptimizer · reddit · 2026-08-31

A user shared an optimized setup for running Qwen3.8-Flash-Next on 4x AMD R9700 GPUs. By using MXFP4-FP8 quantized weights and a custom vLLM Docker image, the setup achieves 120 t/s generation and 12k t/s prefill speeds for a single request. The configuration enables ROCm optimizations, FP8 KV cache, prefix caching, chunked prefill, and MTP speculative decoding, with a full Podman launch command provided.

Original post →

More from Infra

Infra channel →