Trimming MTP draft vocab to 47k boosts DGX Spark code decoding by 21.5% on same hardware
MaziyarPanahi · x · 2026-09-10
A hands-on optimization for Qwen3.8-Flash-Next (NVFP4) on a single NVIDIA DGX Spark: trimming the MTP speculative-decoding draft vocabulary from 248k rows to a code-tuned 47k shrank the draft head from 1.18 to 0.22 GiB and skipped 2.9 GiB per speculative step. Result: single-stream code decoding jumped from 50.6 to 61.5 tok/s (+21.5%), 18% average single-stream gain, 252 tok/s aggregate at 8 streams — with gains decaying as concurrency rises.
More from Infra
- Keras ships ZeroModels: 100+ model families in pure Keras 3, runnable on any backend — fchollet · 2026-09-10
- Cohere moves to NVIDIA Blackwell, cutting token costs and TTFT by 30–50% — cohere · 2026-09-10
- Author uses local AI models to review book manuscripts for $0 in tokens — walkingriver · 2026-09-10
- Musk's Colossus data center fuels massive local backlash, reports Scientific American — scientificamerican · 2026-09-10
- Analyst: Huawei to undercut US AI stack with cheaper chips tuned for Chinese models — 2C_ornot2C · 2026-09-10
- Photon 2.2 Ships Optimized Local Inference for Ampere Through Blackwell GPUs — JFPuget · 2026-09-10