Dual 3090 owners debate adding more cards: bigger local models vs parallel instances
Blues520 · reddit · 2026-09-04
A local LLM hobbyist running dual RTX 3090s (48GB, comfortable with Qwen 27B) asks whether adding two more cards to reach 96GB is worth it.
The trade-off under discussion: stack cards to run larger models (new Qwen Flash, quantized DeepSeek variants) or run two parallel instances to boost coding-workflow throughput. Multi-3090 rig owners weigh in — used 3090s are cheap, but power draw and interconnect efficiency are the real costs.
More from Infra
- Kastard: Edit ComfyUI Workflows Locally, Auto-Sync Models to RunPod GPUs — spacebearbug · 2026-09-04
- Running Qwen3.8 Flash NVFP4 on a single DGX Spark: 1M context at 37 tok/s — BLUECOW009 · 2026-09-04
- Taalas bakes Llama 3.1 8B into custom silicon, hitting 15,000 tokens per second — npew · 2026-09-04
- 146,010 requests in 13 days: OpenRouter endpoints flap wildly, 44 of 241 shift ≥20 points — ringarc · 2026-09-04
- He built a €18k server with 768GB VRAM—now 2-trillion-parameter open models may leave it behind — myreala · 2026-09-04
- DeepSeek plans largest known Huawei chip cluster with 160,000 Ascend processors — The Decoder · 2026-09-04