Running Qwen3 27B agent on 16GB VRAM: MTP draft trades context and quality for speed
Remarkable_Air_8383 · reddit · 2026-10-03
The author serves a local Hermes agent with llama.cpp, running a Qwen3 27B model (iq3s quant) on 16GB VRAM. With MTP drafting enabled, prefill slows down and the model seems less capable, and tight VRAM forces context down to 96k while decode hits 40-60 tps. Disabling MTP restores 128k context but decode drops to 35 tps; the post asks which trade-off to pick.
More from Infra
- Read-only Postgres replicas won't stop LLM agents from bloating your primary DB: 4 failure modes — EmetInteractive · 2026-10-03
- Cathie Wood: 90% of global data center debt financing lands in US at ~8% effective tax — rohanpaul_ai · 2026-10-03
- DeepSeek open-sources DeepGEMM-Ascend: MIT-licensed kernels for Huawei Ascend 950 — lmoroney · 2026-10-03
- DDR4 + 7900XTX Runs Qwen3-Next at 45-50 tok/s via Strata, Double llama.cpp Speed — EmPips · 2026-10-03
- Musk says growth will far exceed forecasts as energy remains the bottleneck — RachelVT42 · 2026-10-03
- DwarfStar 4 (ds4) lets you run DeepSeek V4.1, Qwen and GLM locally — yogthos · 2026-10-03