Can Dual 3060 GPUs Handle 31B Local Inference
MushroomCharacter411 · reddit · 2026-07-12
A user is considering adding a second, cheap **RTX 3060** to their existing local inference machine to fit larger 31B models and caches into VRAM, aiming to escape the current 1.5 t/s or lower speeds. ### Current Setup - i5-8500 - 48GB DDR4 - RTX 3060 12GB (PCIe Gen 3) - Current speed for Gemma 4 26B-A4B Q4 is about 12–15 t/s ### Key Questions - Can a second 3060 make running a **31B dense model at Q4/Q5 quantization** more viable - Will the motherboard's second slot, which only has **x4 bandwidth**, severely bottleneck performance - Can **layer splitting** mitigate the x4 bottleneck The core question of the post is: without spending hundreds of dollars on a larger VRAM GPU, is **dual 3060 + low bandwidth slot** a viable, cost-effective solution for local LLMs?
More from Infra
- Local AI may pay back in 6–7 years and cut long-term costs by 30–40% — DavidLinthicum · 2026-07-21
- TSMC reportedly plans up to 10% chipmaking price hikes in 2027 — kimmonismus · 2026-07-21
- More open models and llama.cpp updates are coming, says Merve Noyan — mervenoyann · 2026-07-21
- Why adding a second LLM provider breaks more than the API surface — Ok_Extension6373 · 2026-07-21
- UK AI datacentres face backlash over heat, noise and land use — nordicinst · 2026-07-21
- Fluidstack raises $830M at $7.5B valuation as Anthropic backs a $50B compute buildout — rohanpaul_ai · 2026-07-21