Dual 5070+5060Ti local LLM setup too slow: is a single RX 7900XTX the fix?
rawdikrik · reddit · 2026-09-29
The author runs local models (STT, memory, a small LLM) on an Unraid server with a 5070 plus a riser-mounted 5060Ti limited to x4. Cross-card inference is bottlenecked at the hardware level, and tools like vllm and ninfer can't fix it. They're considering a single RX 7900XTX for simplicity and ask the community which setup wins for low-stakes local use.
More from Infra
- Buyer told to wait 8-12 months for an H100: the GPU shortage is about power, cooling and talent, not chip supply — ingliguori · 2026-09-29
- Tracking Codex Subs Across Accounts: Switching by Reset Time to Max Out Capacity — idanbeck · 2026-09-29
- Kernel Design Agents optimize Kimi Delta Attention kernels, up to 2.96x speedup — songhan_mit · 2026-09-29
- OpenRouter inference providers slash GLM 5.3 output price from $4.40 to $1.61 in a month — AccBalanced · 2026-09-29
- ASCI Q supercomputer crashed every 6.5 hours until a neutron beam exposed cosmic rays as the culprit — lauriewired · 2026-09-29
- How SNI actually works: the TLS feature powering CDNs and multi-tenant routing — HankYeomans · 2026-09-29