Tuning draft acceptance for Qwen3.6-35B MTP speculative decoding in llama.cpp

Bulky-Priority6824 · reddit · 2026-09-07

A Reddit user asks for help tuning llama.cpp's llama-server running Qwen3.6-35B-A3B (MTP, Q8KXL) with draft-MTP speculative decoding (spec-draft-n-max 5), wondering whether the draft acceptance rate is in the expected range and what else can be tweaked.

The post includes the full launch command with notable settings:

A rare complete reference config for developers running MTP speculative decoding locally.

Original post →

More from Infra

Infra channel →