Dual 5070+5060Ti local LLM setup too slow: is a single RX 7900XTX the fix?

rawdikrik · reddit · 2026-09-29

The author runs local models (STT, memory, a small LLM) on an Unraid server with a 5070 plus a riser-mounted 5060Ti limited to x4. Cross-card inference is bottlenecked at the hardware level, and tools like vllm and ninfer can't fix it. They're considering a single RX 7900XTX for simplicity and ask the community which setup wins for low-stakes local use.

Original post →

More from Infra

Infra channel →