DSpark speculator trained on live SGLang lifts decode throughput 1.89× on B200s
ying11231 · x · 2026-07-23
A team says it trained a DSpark speculator for Inkling NVFP4 using SpecForge against a live SGLang serving engine, instead of an offline dataset.
What changed
- The draft model learned directly from what the target model actually generates in production serving conditions.
- That training setup is meant to preserve accept length at real batch sizes, not just in benchmark scripts.
Reported results
- 1.89× decode throughput over a non-spec baseline at bs=64 on 8×B200, TP8.
- 7–14% faster than Inkling’s built-in MTP at the same batch size.
- 3.66 mean accept length, reaching 4.76 on GSM8K.
Under the hood
- A 5-layer Qwen3-style DFlash parallel-draft backbone
- A rank-256 Markov logit-bias head for intra-block dependencies
- A per-position confidence head to predict acceptance
- 400K online regenerations distilled from the live target engine
The team says the serving command is in the comments for anyone who wants to try it.
Related event: Open-Sourced DSpark Speculator Boosts Inkling Throughput by 1.89x(3 posts)→
More from Infra
- For a 1 GW data center, build 2 GW into the grid — anderssandberg · 2026-07-23
- MCP 2026-07-28 removes sessions and turns the protocol stateless — Substantial-Heat-321 · 2026-07-23
- Google is spending $200B+ on cloud and compute, Beff Jezos says — beffjezos · 2026-07-23
- Alibaba Cloud says its Zhenwu M890 supernode now runs Qwen3.8 for inference — zephyr_z9 · 2026-07-23
- Alibaba listing shows a $5,000 container data center in the shopping cart — bronzeagepapi · 2026-07-23
- 20VC maps the AI market debates around Kimi, OpenRouter, Fireworks and Nvidia — 20VC · 2026-07-23