Running a 27B Qwen model on RTX 3060: full llama.cpp config hits 10-20 tok/s

SummarizedAnu · reddit · 2026-09-10

A Reddit user shares a complete, working llama.cpp setup for running a Qwen3.8-27B IQ3XXS quant (GSQ-RCO-MTP GGUF) on a 12GB RTX 3060:

A concrete reference for anyone running 27B-class models locally on consumer GPUs.

Original post →

More from Infra

Infra channel →