Quad R9700 with vLLM-Radiance hits 17.6k tok/s prefill on EPYC

im_EDEN · reddit · 2026-09-03

A Redditor benchmarked four Radeon AI Pro R9700s running vLLM-Radiance: official vLLM was painfully slow, but the Radiance fork reached 17,636 tok/s prefill and 36.6 tok/s decode, peaking at 106 tok/s with an 80% MTP acceptance rate. Setup: Gigabyte MZ32-AR0 + EPYC 7282, Qwen 3.8 27B fp8 at 262k context, one card on PCIe 4x8. Notably all four cards run at 100% utilization even though only dual-GPU setups are officially supported.

Original post →

More from Infra

Infra channel →