Dual GPU inference with RTX 4090 + 3060: speed impact and optimization tips
cosmoschtroumpf · reddit · 2026-07-30
A user asks about adding a 3060 to a 4090 for dual GPU inference. Community discusses speed reduction due to PCIe bandwidth and inter-GPU communication, but increased context length. Recommends llama.cpp or vLLM with memory/compute splitting. Suggests trying Q4 quantization for balance.
More from Infra
- OpenAI Says GPT-5.6 Sol Self-Optimizes: 20% Lower Serving Costs — OpenAI · 2026-07-30
- Microsoft and Meta Show AI Infrastructure Spending Pays Off: Azure Grows 43%, Copilot Hits 30M Paid Seats — luisdans · 2026-07-30
- Local AI Boom Could Make Storage Drive Manufacturers a Fortune — cocktailpeanut · 2026-07-30
- Meta Shares Plunge 9% After-Hours as AI Spending Crushes Margins & Cash Flow — ivan_bezdomny · 2026-07-30
- US Commerce Dept Allocates $874M to Accelerate Semiconductor R&D — imjustnewatai · 2026-07-30
- Valar Atomics Founder: Cheap Energy Will Always Create Its Own AI Demand — No Priors · 2026-07-30