Can Dual 3060 GPUs Handle 31B Local Inference

MushroomCharacter411 · reddit · 2026-07-12

A user is considering adding a second, cheap **RTX 3060** to their existing local inference machine to fit larger 31B models and caches into VRAM, aiming to escape the current 1.5 t/s or lower speeds. ### Current Setup - i5-8500 - 48GB DDR4 - RTX 3060 12GB (PCIe Gen 3) - Current speed for Gemma 4 26B-A4B Q4 is about 12–15 t/s ### Key Questions - Can a second 3060 make running a **31B dense model at Q4/Q5 quantization** more viable - Will the motherboard's second slot, which only has **x4 bandwidth**, severely bottleneck performance - Can **layer splitting** mitigate the x4 bottleneck The core question of the post is: without spending hundreds of dollars on a larger VRAM GPU, is **dual 3060 + low bandwidth slot** a viable, cost-effective solution for local LLMs?

Original post →

More from Infra

Infra channel →