Trimming MTP draft vocab to 47k boosts DGX Spark code decoding by 21.5% on same hardware

MaziyarPanahi · x · 2026-09-10

A hands-on optimization for Qwen3.8-Flash-Next (NVFP4) on a single NVIDIA DGX Spark: trimming the MTP speculative-decoding draft vocabulary from 248k rows to a code-tuned 47k shrank the draft head from 1.18 to 0.22 GiB and skipped 2.9 GiB per speculative step. Result: single-stream code decoding jumped from 50.6 to 61.5 tok/s (+21.5%), 18% average single-stream gain, 252 tok/s aggregate at 8 streams — with gains decaying as concurrency rises.

Original post →

More from Infra

Infra channel →