Running Qwen3.8-27B on 2x3090: 200K Context with F16 KV, Vision, and Reasoning
Sisuuu · reddit · 2026-08-15
A user successfully deployed the Q8KXL quantization of Qwen3.8-27B on a dual RTX 3090 setup, achieving a 200K context window with F16 KV cache and vision support enabled. The post details the memory optimization benefits of the model's hybrid architecture (16 full attention layers + 48 linear attention layers) and the specific llama.cpp configuration required to avoid OOM errors. It provides the full command line recipe for enabling speculative decoding and reasoning features, reporting final speeds of 73 tok/s for code and 58 tok/s for prose.
More from Infra
- The 2026 Mastering Databricks Roadmap for Data Engineers — Zachly · 2026-08-15
- Efficiency, not raw model performance, is the new AI infra battlefield — rudina11 · 2026-08-15
- AMD proposes 'threads per megawatt' as a new metric for the agentic AI era — xiaosun86 · 2026-08-15
- Can AI-Generated Code Scale? Cosmos DB Demo Tests Agent Performance Under Load — adnan_hashmi · 2026-08-15
- Open Source LFM 2.5-2.6B Model Crashes Under 4.3B Token Load — maximelabonne · 2026-08-15
- Huawei adopts HBF and other techniques to mitigate HBM shortage — bookwormengr · 2026-08-15