Running Qwen3.8-27B on 2x3090: 200K Context with F16 KV, Vision, and Reasoning

Sisuuu · reddit · 2026-08-15

A user successfully deployed the Q8KXL quantization of Qwen3.8-27B on a dual RTX 3090 setup, achieving a 200K context window with F16 KV cache and vision support enabled. The post details the memory optimization benefits of the model's hybrid architecture (16 full attention layers + 48 linear attention layers) and the specific llama.cpp configuration required to avoid OOM errors. It provides the full command line recipe for enabling speculative decoding and reasoning features, reporting final speeds of 73 tok/s for code and 58 tok/s for prose.

Original post →

More from Infra

Infra channel →