Qwen 27B on 2× RX 7900 XT: 66.5 TPS single-stream, still short of claimed 100+
EqualCryptographer67 · reddit · 2026-10-06
A Reddit user benchmarked Qwen 27B and Flash Next local inference on 2× RX 7900 XT (20GB each) with 128GB DDR4, publishing a detailed spreadsheet of results:
- Qwen3.8-27B IQ4XS (ROCm + MTP3): 46.4 TPS generation; switching to IQ2XXS quantization reached 66.5 TPS
- Flash Next UD-IQ4XS (93.7GB): only 8.6 TPS on one GPU via ROCm, and worse at 5.5–6.5 TPS with dual-GPU Vulkan tensor split 1:1
- Two parallel IQ4 copies at 16 concurrent requests hit 103.5 TPS combined throughput, but single-response 100 TPS was never achieved
- Gotchas: tensor split across cards produced broken text until --no-mmap fixed it; the PowerColor card sits on a PCIe 3.0 x1 slot, a likely bottleneck
The author questions whether reported 100+ TPS figures on 16GB VRAM refer to single responses or aggregated concurrent throughput, and asks for advice on CPU/RAM offloading, the x1 link, or backend settings. Test conditions: 8k context, f16 KV, thinking off.
More from Infra
- Can a 128GB M5 Max Mac Studio Handle Concurrent Local LLM Agents? — Simple_Telephone_867 · 2026-10-06
- MiniMax discloses 70-80% inference margins, fueling debate on AI subscription subsidies — menhguin · 2026-10-06
- ~90% of frontier lab compute now goes to post-training and inference — IanAndrewsDC · 2026-10-06
- KLIF open-sources a Rust front-end that manages llama.cpp, vLLM and TTS servers in one window — Koksny · 2026-10-06
- Dev burns 842B tokens in September — $409k at API list price, pays just 3.4% via subscription — doodlestein · 2026-10-06
- One Dot burns 1.6B tokens/day on a $100 subscription — roughly $540k/month in API-equivalent compute — DarthSilent · 2026-10-06