Dual R9700 AI Pro hits only ~50 t/s on Qwen3-27B-FP8 with vLLM speculative decoding
YehowaH · reddit · 2026-08-24
A user deployed official Qwen3-27B-FP8 with MTP=3 speculative decoding on two AMD R9700 AI Pro GPUs (TP=2, PCIe Gen5 x8, Ryzen 5 7400, DDR5-6000) via vLLM. Real generation throughput is only 50 t/s with 30 t/s accepted, mean acceptance length 2.7-2.8, and 60% average draft acceptance rate. The setup uses the andysalerno/r9700-serving repo (unified aiter attention, ROCm 7.14, latest vLLM/FlashAttention/aiter). They asked the community whether these numbers are reasonable and where bottlenecks might lie.
More from Infra
- Robot Boom Sparks Supply Chain Growth: Motors and Drives See Cost Drops — CyberRobooo · 2026-08-24
- Agentic Payment Protocols: x402 Leads, Stripe Enters with MPP — MountainAssignment36 · 2026-08-24
- Google Gemma-4-26B Ported to Apple Silicon with Half Memory Footprint — jasonkneen · 2026-08-24
- Sovereign AI market valued at $1.5T as reliance on US Big Tech falls — mikeflache · 2026-08-24
- Ask HN: Best way to add vision support to Deepseek v4 Flash on DGX? — StartupTim · 2026-08-24
- Solutions for Running Agents on Macs: Prevent Sleep and Cloud Fleet Management — tedddyoweh · 2026-08-24