Dual 7900XTX setup gets only 11 tps on Qwen 3.8 Next, slower than single 3090s
Nota_ReAlperson · reddit · 2026-09-05
A Reddit user running Qwen 3.8 Next on the latest llama.cpp reports only 11 tokens per second with two 7900XTX GPUs and 128GB DDR4 — worse than what others report on single RTX 3090s or 9700s — and asks for troubleshooting suggestions.
More from Infra
- SGLang's Breakable CUDA Graph speeds prefill graph building by 3.8–5.2x — ying11231 · 2026-09-05
- SemiAnalysis: OpenAI's ASIC program is leverage — Altman wins even if the chip loses — MarvinTBaumann · 2026-09-05
- 2027 will be peak year of AI compute constraint; relief arrives in 2028, analyst argues — BenBajarin · 2026-09-05
- Dev builds P2P network to seed open-weight models, fearing a NVIDIA-Hugging Face deal — Relevant-Magic-Card · 2026-09-05
- NVIDIA lays out five practical guidelines for speculative decoding to speed up LLM inference — NVIDIAAI · 2026-09-05
- NVIDIA Details Hardware-Friendly LLM Design and Five Speculative Decoding Guidelines — NVIDIAAI · 2026-09-05