Qwen3.8-27B coder quant fits a 24GB GPU with 262k context at 40 t/s

W61k3r · reddit · 2026-10-02

A community quant packs Qwen3.8-27B CODER into a single 24 GB GPU with 262k context: an IQ4XS imatrix GGUF (18.35 GB) with the MTP draft head kept at Q80, so speculative decoding works without a separate draft model. It was quantized and 'abliterated' by a tool called LexiPanel, whose fit planner chose the tensor mix to fill the card.

Real-world numbers on an RX 7900 XTX (llama.cpp Vulkan): median 41.2 t/s decode at 65k–131k context and 36.0 t/s at 131k–171k; MTP draft acceptance median 85% (2.7 tokens per step); prefill 387 t/s on a cold 108k prompt; 24.2/24.6 GB VRAM in use at 245k tokens with a q41 KV cache and vision projector on CPU. The author runs it daily for agentic coding and notes it 'doesn't get dumber while coding' like other fine-tunes.

Full llama-server command is provided, including --spec-type draft-mtp, Qwen-recommended sampling (temp 1.0, top-p 0.95, top-k 20), and tips like --cache-ram for agent loops.

Original post →

More from Infra

Infra channel →