Quad R9700 with vLLM-Radiance hits 17.6k tok/s prefill on EPYC
im_EDEN · reddit · 2026-09-03
A Redditor benchmarked four Radeon AI Pro R9700s running vLLM-Radiance: official vLLM was painfully slow, but the Radiance fork reached 17,636 tok/s prefill and 36.6 tok/s decode, peaking at 106 tok/s with an 80% MTP acceptance rate. Setup: Gigabyte MZ32-AR0 + EPYC 7282, Qwen 3.8 27B fp8 at 262k context, one card on PCIe 4x8. Notably all four cards run at 100% utilization even though only dual-GPU setups are officially supported.
More from Infra
- Inference Engineering Is Just a Recipe: vLLM/SGLang, Replicas, Cache-Aware Routing — GabGarrett · 2026-09-03
- Databricks pitches agent-native data infrastructure, Lakebase Postgres at VLDB 2026 — matei_zaharia · 2026-09-03
- Lablup, Maker of GPU Orchestrator Backend.AI, Joins PyTorch Foundation as Silver Member — PyTorch · 2026-09-03
- Cursor cloud agents can now run on your own infrastructure, Mac Minis included — mattyp · 2026-09-03
- VideoDeltaNet open-sources hybrid attention that speeds up MiniMax H3 video generation up to 90x — realmrfakename · 2026-09-03
- Analyst: NVIDIA Could Become Intel Foundry's 'Customer Zero' as a Second Source Beyond TSMC — BenBajarin · 2026-09-03