Is Running a 27B Model on a 3090 Plus a 3070 Worth It?
drone_syndrome · reddit · 2026-07-11
The author ran Qwen3.6-27B (Q5KS, llama.cpp fork + speculative decoding) on a single 3090 24GB. At a 140K context, VRAM was nearly maxed out, yet it still achieved about 100 tk/s for code generation and 45 tk/s for prose.
They are now considering pairing it with a 3070 8GB and wondering if the combo would offer real benefits or just bottleneck the system due to the 3070's slower bandwidth:
- Would the extra 8GB allow for larger quantization (like Q6) or longer context?
- Would a 3090+3070 setup run a 27B dense model faster than a single 3090?
- Is a 660W power supply viable for a dual-GPU setup with power limits?
- Or would it be better to sell the 3070 and save up for a second 3090?
More from Infra
- China’s AI arms race is increasingly defined by chips, data centers, and open models — BenBajarin · 2026-07-22
- Agent search bottlenecks are now about variance, not raw latency — rohanpaul_ai · 2026-07-22
- Gavin Baker argues Nvidia may be one of open source AI’s biggest supporters — GavinSBaker · 2026-07-22
- AI Power Demand Exposes US Energy Gap, Urging Shift from Scarcity to Abundance — bradneuberg · 2026-07-22
- Gavin Baker says Nvidia’s $630B figure would be system revenue, not all Nvidia’s — GavinSBaker · 2026-07-22
- A Firecracker-based platform says it can host 6,000 AI agents on one 256 GB server — maritime_sh · 2026-07-22