Running Qwen 27B Locally on 2× RTX 5070 Ti: A Cost-Effective Inference Setup

val_in_tech · reddit · 2026-08-06

The author shares a cost-effective local LLM inference setup, running the Qwen 3.6 27B model using two RTX 5070 Ti GPUs. A single 5070 Ti provides 896 GB/s of memory bandwidth, making it highly suitable for bandwidth-bound dense models.

Performance metrics:

The author suggests that for users accustomed to RTX 3090-level performance, the RTX 5070 Ti is a powerful option that supports all modern features at a relatively affordable price point.

Original post →

More from Infra

Infra channel →