Qwen on 4xR9700: 120 t/s Gen and 12k t/s Prefill with Optimized vLLM
sloptimizer · reddit · 2026-08-31
A user shared an optimized setup for running Qwen3.8-Flash-Next on 4x AMD R9700 GPUs. By using MXFP4-FP8 quantized weights and a custom vLLM Docker image, the setup achieves 120 t/s generation and 12k t/s prefill speeds for a single request. The configuration enables ROCm optimizations, FP8 KV cache, prefix caching, chunked prefill, and MTP speculative decoding, with a full Podman launch command provided.
More from Infra
- SK Hynix breaks ground on Indiana HBM plant, targeting HBM4e mass production in 2029 — Beth_Kindig · 2026-08-31
- Qwen 3.8 Flash Next runs at 3.5 tok/s on mid-range Android phone — dai_app · 2026-08-31
- The Boring Company's sales dilemma: 20x cost advantage but no sales team — PTrubey · 2026-08-31
- Running 182B Qwen on 4080: 8 tok/s via SSD offloading — Desperate-Data-3747 · 2026-08-31
- AI is transforming how energy systems are monitored, maintained and operated — ingliguori · 2026-08-31
- Uber Cuts AI Costs 52% While 10xing Usage via Agentic Workflow — alvelda · 2026-08-31