Running Qwen3.8-27B on dual RTX 3060s: 50 tok/s recipe

Ecstatic-Wash-7667 · reddit · 2026-08-15

A Reddit user shares a llama.cpp configuration for running Qwen3.8-27B on dual RTX 3060 12GB, achieving 50.6 tok/s generation with MTP speculative decoding, 573 tok/s prefill, and 23.2/24.0 GiB VRAM usage.

Related event: Qwen3.8-27B Hits 50 tok/s on Dual RTX 3060s(2 posts)→

Original post →

More from coding & agent

coding & agent channel →