Qwen3.8-27B Hits 50 tok/s on Dual RTX 3060s

A Reddit user demonstrated running Qwen3.8-27B on dual RTX 3060 12GB GPUs, achieving 50.6 tok/s using llama.cpp and MTP speculative decoding.

2026-08-15 ~ 2026-08-15 · 2 related posts