Is Running a 27B Model on a 3090 Plus a 3070 Worth It?

drone_syndrome · reddit · 2026-07-11

The author ran Qwen3.6-27B (Q5KS, llama.cpp fork + speculative decoding) on a single 3090 24GB. At a 140K context, VRAM was nearly maxed out, yet it still achieved about 100 tk/s for code generation and 45 tk/s for prose.

They are now considering pairing it with a 3070 8GB and wondering if the combo would offer real benefits or just bottleneck the system due to the 3070's slower bandwidth:

Original post →

More from Infra

Infra channel →