2x AMD R9700 beats 3x running Qwen 3.8 Flash Next locally at 35 t/s
MarcusAurelius68 · reddit · 2026-09-05
A Reddit user benchmarked Qwen 3.8 Flash Next on AMD R9700 32GB GPUs and found that a 2-GPU setup actually beats 3 GPUs. The culprit: in the 3-GPU config on an X570 board, the third card sits on a chipset x4 slow lane that drags down overall performance.
Key details:
- Hardware: X570 at x8/x8, Ryzen 9 5900XT, 64GB DDR4, Windows 11.
- Running the AD-4.27bpw Q4KM quant (33-part GGUF, 131k context) with tuning (no MTP yet), the 2-GPU setup hit 242 t/s prompt processing and 35 t/s generation.
- The full llama-cli command is shared, including --fa on, q80 KV cache quantization, --jinja, and sampling settings.
A directly reproducible reference for anyone running large-context local models on consumer hardware.
More from Infra
- There's no agreed way to value a GPU running inference—and compute futures now settle on these indexes — AccBalanced · 2026-09-05
- AI Now on data center boom: community pushback and 'they won't build them where they live' — AINowInstitute · 2026-09-05
- Gemma 4 Runs 151.4% Faster on Mac via Community MLX Inference Optimization — gajesh · 2026-09-05
- SGLang's Breakable CUDA Graph speeds prefill graph building by 3.8–5.2x — ying11231 · 2026-09-05
- SemiAnalysis: OpenAI's ASIC program is leverage — Altman wins even if the chip loses — MarvinTBaumann · 2026-09-05
- 2027 will be peak year of AI compute constraint; relief arrives in 2028, analyst argues — BenBajarin · 2026-09-05