Dual R9700 AI Pro hits only ~50 t/s on Qwen3-27B-FP8 with vLLM speculative decoding

YehowaH · reddit · 2026-08-24

A user deployed official Qwen3-27B-FP8 with MTP=3 speculative decoding on two AMD R9700 AI Pro GPUs (TP=2, PCIe Gen5 x8, Ryzen 5 7400, DDR5-6000) via vLLM. Real generation throughput is only 50 t/s with 30 t/s accepted, mean acceptance length 2.7-2.8, and 60% average draft acceptance rate. The setup uses the andysalerno/r9700-serving repo (unified aiter attention, ROCm 7.14, latest vLLM/FlashAttention/aiter). They asked the community whether these numbers are reasonable and where bottlenecks might lie.

Original post →

More from Infra

Infra channel →