Qwen3.8 Flash hits 74 tok/s single-stream, 212 tok/s aggregate on one DGX Spark — open vLLM recipe
DimeRhyme · reddit · 2026-09-29
Results
The author spent weeks tuning a vLLM serving recipe for Qwen3.8 Flash on a single DGX Spark (GB10), fully open-sourced with raw benchmark data:
- 74 tok/s peak single-stream (typical setups land at 35-45), 60-70 on normal requests, and 212 tok/s across 8 streams.
- Cold prefill 2-3.4x faster than the forked recipe (16K prompt: 4,016 vs 1,171 tok/s).
- Full 262K context; cached coding prompts start replying in 0.57s; quality matches the original within noise (93.1%/93.3% on a 492-question suite, top-1 agreement shift of 0.06 pts).
What made the difference
- Running the model's own MTP head densely: 3.7 tokens per verify step — the biggest lever, with wrong guesses replaced by the full model so quality is unchanged.
- Pruning the draft head vocab from 248K to 65K ids: the draft pass is memory-bound on GB10; the target model still verifies against the full vocab.
- Quantization: W4A16 AutoRound for MoE experts, FP8 for side layers, INT8 lmhead — deliberately avoiding 3-bit/NVFP4 so speed comes from the serving path.
- GB10-tuned low-latency GEMM for decode-time matmuls plus sort-free top-k in verification.
- A faster per-layer embedding gather path removing serial page faults during prefill.
- Decode step went from 68.3ms to 52.3ms on an agent-shaped coding workload.
- A failed experiment: doubling prefill chunk to 16,384 tokens gave no gain and strained memory.
Footprint
71 GiB model, 16 GB KV pool, 16 GiB free under load; GPU at 35-37W median while generating. Most tricks apply to any MTP/speculative-decoding setup.
More from Infra
- Open-weight 0.8B/2B System 1 decision models match Jev at 83.1%, trained fully on local hardware — Usual_Maximum7673 · 2026-09-29
- Redditor predicts sub-$1000 device running SOTA models will spawn the next big company — Robert__Sinclair · 2026-09-29
- Only 3 of ~6,000 data center projects hit by AI buildout moratoriums: SemiAnalysis — MatthewBerman · 2026-09-29
- Cloudflare birthday week ships 8 open source updates: forge, vinext 1.0, native Rust in Workers — ritakozlov · 2026-09-29
- Google Trends' #1 US region for every query is tiny Cheyenne, Wyoming — likely bot traffic — lilyraynyc · 2026-09-29
- Sentdex has run 4B+ tokens locally on GLM 5.3 Flash — his most-used local model ever — Sentdex · 2026-09-29