2× R9700 local rig runs Qwen3.8-27B at 111 tok/s for ~€4k, €1k under a 5090
smallDeltaBigEffect · reddit · 2026-09-07
A Reddit user shares full specs and first-day benchmarks of a local inference rig with 2× AMD Radeon AI PRO R9700 (32GB each) — total cost €4,000, over €1k less than a single RTX 5090.
- Build: Ryzen 7500F, 64GB DDR5 6400, Asus ProArt X870E, both cards at PCIe 5.0 x8; one card runs hot, planned power-limit to 210W plus undervolting
- Qwen3.8-27B (Quark AWQ MXFP4, vLLM Radiance TP2): 111.4 tok/s single-stream decode median, ITL 1% low 77.9 tok/s, TTFT 81ms, 131k context, 4,224 tok/s prefill at 2k prompts
- Native FP8 weights: 87.6 tok/s, 16k context, 8 concurrent sequences
- Qwen3.8-Flash-Next (UD-IQ4XS GGUF, R9V fork tiered expert offload): 35.4 tok/s; surprisingly, SATA SSD offload wasn't terrible
- Use cases: deep research, summarization, image generation, light coding; next up are context degradation and KV quants
Solid reference data for cost-effective dual-AMD vLLM local deployment.
More from Infra
- AI meme: turning off the tap while brushing teeth to "conserve water for the datacenter buildout" — EigenGender · 2026-09-07
- AMD MI355X beats NVIDIA B300 on tokens-per-dollar in AgentX — AccBalanced · 2026-09-07
- RandKV ships as pip-installable random KV-cache eviction for Transformers, reports honest negative perf results — atease01 · 2026-09-07
- Offline village AI: $5,000 budget to build a local LLM machine for basic Q&A, seeking GPU advice — Potential_Low_1183 · 2026-09-07
- Custom llama.cpp Branch Adds Expert Expansion for MoE Models — Specific-Tax-6700 · 2026-09-07
- Developer earns just $2.60 per cycle running a bot on OpenAI-subsidized tokens — TheMoonMidas · 2026-09-07