Open-Sourced DSpark Speculator: Boosts Inkling Decode Throughput by 1.89X
BanghuaZ · x · 2026-07-23
A developer has trained a new DSpark speculator for the Inkling NVFP4 model, built end-to-end with SpecForge on a live SGLang target engine, significantly improving inference efficiency.
Performance metrics include:
- 1.89x decode throughput improvement over non-speculative decoding at bs=64 (8xB200, TP8).
- 7-14% faster than Inkling's built-in MTP at the same batch size.
- Mean accept length of 3.66, reaching up to 4.76 on the GSM8K dataset.
Under the hood, it utilizes a 5-layer Qwen3-style DFlash parallel-draft backbone, a Rank-256 Markov logit-bias head, a per-position confidence head for acceptance prediction, and was distilled online from 400K Inkling regenerations.
More from Infra
- AI job listings mentioning evals rise to 10.7% by July 2026 — HamelHusain · 2026-07-23
- Writer Study: Optimizing AI Harness Reduces Costs by 41% Without Losing Accuracy — bendee983 · 2026-07-23
- Voice assistant tool calls sped up instantly after moving the backend to Europe — ur_piyo_a_hoe · 2026-07-23
- Meta overhauls blob storage to cut GPU stalls across exabyte-scale clusters — Meta_Engineers · 2026-07-23
- Claude Code free credits quietly flipped one user’s spend limit to unlimited — StreamfireEU · 2026-07-23
- Samsung Distributor's HBM Shipments to China Spike Post-Export Controls — ohlennart · 2026-07-23