Qwen 27B on a 4060 Ti at 6-7 t/s: What Else Can Squeeze Local Inference Speed?
thatoneshadowclone · reddit · 2026-09-03
Running Qwen 27B (IQ3S Unsloth) as a coding agent on a 4060 Ti 16GB with llama.cpp, the author gets 6-7 t/s on reasoning and 11 t/s on code generation (1-1.5h per task). Current setup: dual GGUF with draft-dflash speculative decoding (max 3 drafts), -nkvo, 128K context, -fa on, q80 KV cache, b2048/ub512, full GPU offload. Asking what else can be tuned given the priority on large context.
More from Infra
- Reliance Jio launches JioPC cloud PC across India starting at ₹11/day — NirantK · 2026-09-03
- Arm CEO Rene Haas on chips, Meta's Arm AGI CPU, SoftBank and robotics — No Priors · 2026-09-03
- ARBR open-sources an MIT-licensed routing and governance layer for multi-provider LLM traffic — Genie-Tickle-007 · 2026-09-03
- Arm CEO Rene Haas: from IP licensing to building chips, including an AGI CPU for Meta — No Priors · 2026-09-03
- ClickHouse CEO: agents have no personas, latency is the key requirement — bibryam · 2026-09-03
- DeepEP + RoCE Doesn't Work, Warns Engineer Rerunning Benchmarks — TheZachMueller · 2026-09-03