Dual RTX 3060 Local LLM Setup: 100k Context at 600 tok/s, $2k Upgrade Paths
gnoremepls · reddit · 2026-09-19
A Reddit user shares real numbers from a dual second-hand RTX 3060 (2×12GB, $300 each) setup running Qwen3.8:27B: 100k context at Q4, 600 tok/s input and 45 tok/s output, plus idle 128GB DDR5-5600. With a $2k budget they weigh adding two more 3060s in a 4-GPU machine (PCIe lane bottleneck concerns), a single RTX 3090 ($2000 for only +8GB), or dual 24GB Intel B60 cards ($800 each), with community advice on offloading and scaling.
More from Infra
- Reading a Pretraining Run: A P0/P1/P2 Metric System for Monitoring LLM Pretraining — SonglinYang4 · 2026-09-19
- Dual RTX 5060 Ti Only Gets 10 t/s on Qwen3.8-Flash-Next, Seeking Config Advice — MkGod · 2026-09-19
- Hacking open-source model behavior with sglang's scoring endpoint, no fine-tuning needed — BLUECOW009 · 2026-09-19
- Distilling DeepSeek V4 Flash to a 4B model on DGX Spark: 26 hours, 22ms per judgment — Dan_Jeffries1 · 2026-09-19
- Positron raises $875M at $5B valuation as its co-founder calls anti-data-center talk a "Chinese psyop" — 20VC · 2026-09-19
- GitHub Next open-sources LocalJev, a local Jev-compatible API built on oMLX and DiffusionGemma — gaganghotra_ · 2026-09-19