Dual RTX 3090 running Qwen3.8-27B locally at 50-65 tok/s — is that normal?
sugarfreecaffeine · reddit · 2026-08-20
A Reddit user shares their setup running Qwen3.8-27B (unsloth dynamic Q6KXL quant) locally with llama.cpp on Windows 10, dual RTX 3090 (48GB VRAM), for local agentic coding. They see 50-65 tokens/sec generation, sometimes dipping to the 40s, and ask if that's expected. The full llama-server command is included — 262K context, q80 KV cache, flash attention, MTP speculative decoding — a useful reference for local deployment tuning.
More from coding & agent
- Max Drake: Canvas as Boundary Object Between You and Your Agent, Sketch to Code — max__drake · 2026-08-20
- Context Window Budgeting: Making every token count — blaizedsouza · 2026-08-20
- AI as a library, not just memory — dfinke · 2026-08-20
- LLM is just one part of a coding agent; the harness is the real system — blaizedsouza · 2026-08-20
- LATAM Airlines Cuts Agent Out-of-Scope Rate to 1% Using Production Traces — LangChain · 2026-08-20
- Grok Bot Demo: Learned Expenses by Watching Once, Automates Backlog Without Scripts — eyishazyer · 2026-08-20