Qwen3.8-27B coder quant fits a 24GB GPU with 262k context at 40 t/s
W61k3r · reddit · 2026-10-02
A community quant packs Qwen3.8-27B CODER into a single 24 GB GPU with 262k context: an IQ4XS imatrix GGUF (18.35 GB) with the MTP draft head kept at Q80, so speculative decoding works without a separate draft model. It was quantized and 'abliterated' by a tool called LexiPanel, whose fit planner chose the tensor mix to fill the card.
Real-world numbers on an RX 7900 XTX (llama.cpp Vulkan): median 41.2 t/s decode at 65k–131k context and 36.0 t/s at 131k–171k; MTP draft acceptance median 85% (2.7 tokens per step); prefill 387 t/s on a cold 108k prompt; 24.2/24.6 GB VRAM in use at 245k tokens with a q41 KV cache and vision projector on CPU. The author runs it daily for agentic coding and notes it 'doesn't get dumber while coding' like other fine-tunes.
Full llama-server command is provided, including --spec-type draft-mtp, Qwen-recommended sampling (temp 1.0, top-p 0.95, top-k 20), and tips like --cache-ram for agent loops.
More from Infra
- llama.cpp adds decision models: /v1/systemone scores options in a single forward pass — ggerganov · 2026-10-02
- Deutsche Bank Initiates FormFactor at Buy With $200 Target on Nvidia GPU Probe Card Share — demian_ai · 2026-10-02
- Why Memory Prices Keep Climbing Amid AI Demand — FinanceYF5 · 2026-10-02
- Jensen Huang on power constraints: tokens/per-watt economics favor Nvidia as 55 of 80 cloud partners sit outside the US — BenBajarin · 2026-10-02
- llama.cpp adds Decision Models, expanding local inference capabilities — paf1138 · 2026-10-02
- Agents aren't GPU-bound: tool execution, memory bandwidth and sandbox overhead are the real bottleneck — ai · 2026-10-02