Running Qwen 27B on RTX 3060+2060 Yields Only 5-6 TPS
sheriffoftiltover · reddit · 2026-08-26
A user ran Qwen3.8-27B (Q4KL) on an RTX 3060 + RTX 2060 setup with 52GB DDR4 RAM, getting only 5-6 tokens/s. Memory headroom is too small for reasonable context lengths, making it fun for experiments but impractical. They ask whether specialized engines or optimizations exist for older GPUs and low-VRAM setups, mentioning ninfer discussions around the RTX 5090.
More from Infra
- NASA seeks Starlink for real-time data at 50,000 feet — XFreeze · 2026-08-26
- Chip benchmarks show leading perf/Watt and perf/$, highlighting Agent vs. Chat workloads — AccBalanced · 2026-08-26
- Applied Compute Launches AC2 Agent Cloud for Training and Serving Custom Models — rhythmrg · 2026-08-26
- Leasing an M5 Ultra Mac Studio costs slightly more than a Claude Max sub — Hesamation · 2026-08-26
- OpenAI, Google, Amazon custom chips threaten Nvidia ecosystem dominance — bindureddy · 2026-08-26
- AI Gateway Patterns: Managing heterogeneity and routing flexibility — philipkiely · 2026-08-26