SGLang DSpark reaches 1.89× decode throughput on 8×B200
BanghuaZ · x · 2026-07-23
A DSpark speculator for Inkling NVFP4 was trained end-to-end with SpecForge on a live SGLang target engine.
- 1.89× decode throughput vs. a non-speculative baseline at batch size 64 on 8×B200 with TP8
- 7–14% faster than Inkling’s built-in MTP at the same batch size
- 3.66 average accept length, reaching 4.76 on GSM8K
Implementation details:
- A 5-layer Qwen3-style DFlash parallel-draft backbone
- A rank-256 Markov logit-bias head for intra-block dependencies
- A per-position confidence head to predict acceptance
- Distillation from 400K Inkling regenerations
The post invites readers to try it, with the serving command shared in the comments.
Related event: Open-Sourced DSpark Speculator Boosts Inkling Throughput by 1.89x(3 posts)→
More from Infra
- Multi-agent systems fail less on reasoning than on orchestration, says one production team — njanChe1 · 2026-07-23
- AI video dubbing costs about $5–7 per finished minute once lip sync is included — Madmahi25 · 2026-07-23
- Nebius shows its first NVIDIA Vera Rubin NVL72 rack in Finland — demian_ai · 2026-07-23
- A Bittensor subnet launches inference at roughly half the usual price — markjeffrey · 2026-07-23
- Ascend SuperPOD optimization lifts DeepSeek-V4 post-training MFU to 34.22% — pmttyji · 2026-07-23
- Turbopuffer halves queue time after fixing deceptively hard autoscaling — DanielLockyer · 2026-07-23