Running Qwen3 27B agent on 16GB VRAM: MTP draft trades context and quality for speed

Remarkable_Air_8383 · reddit · 2026-10-03

The author serves a local Hermes agent with llama.cpp, running a Qwen3 27B model (iq3s quant) on 16GB VRAM. With MTP drafting enabled, prefill slows down and the model seems less capable, and tight VRAM forces context down to 96k while decode hits 40-60 tps. Disabling MTP restores 128k context but decode drops to 35 tps; the post asks which trade-off to pick.

Original post →

More from Infra

Infra channel →