Running Qwen 27B on RTX 3060+2060 Yields Only 5-6 TPS

sheriffoftiltover · reddit · 2026-08-26

A user ran Qwen3.8-27B (Q4KL) on an RTX 3060 + RTX 2060 setup with 52GB DDR4 RAM, getting only 5-6 tokens/s. Memory headroom is too small for reasonable context lengths, making it fun for experiments but impractical. They ask whether specialized engines or optimizations exist for older GPUs and low-VRAM setups, mentioning ninfer discussions around the RTX 5090.

Original post →

More from Infra

Infra channel →