Dual RTX 3090 running Qwen3.8-27B locally at 50-65 tok/s — is that normal?

sugarfreecaffeine · reddit · 2026-08-20

A Reddit user shares their setup running Qwen3.8-27B (unsloth dynamic Q6KXL quant) locally with llama.cpp on Windows 10, dual RTX 3090 (48GB VRAM), for local agentic coding. They see 50-65 tokens/sec generation, sometimes dipping to the 40s, and ask if that's expected. The full llama-server command is included — 262K context, q80 KV cache, flash attention, MTP speculative decoding — a useful reference for local deployment tuning.

Original post →

More from coding & agent

coding & agent channel →