SGLang DSpark reaches 1.89× decode throughput on 8×B200
BanghuaZ · x · 2026-07-23
A DSpark speculator for Inkling NVFP4 was trained end-to-end with SpecForge on a live SGLang target engine.
- 1.89× decode throughput vs. a non-speculative baseline at batch size 64 on 8×B200 with TP8
- 7–14% faster than Inkling’s built-in MTP at the same batch size
- 3.66 average accept length, reaching 4.76 on GSM8K
Implementation details:
- A 5-layer Qwen3-style DFlash parallel-draft backbone
- A rank-256 Markov logit-bias head for intra-block dependencies
- A per-position confidence head to predict acceptance
- Distillation from 400K Inkling regenerations
The post invites readers to try it, with the serving command shared in the comments.
Related event: Open-Sourced DSpark Speculator Boosts Inkling Throughput by 1.89x(3 posts)→
More from Infra
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11