Open-Sourced DSpark Speculator: Boosts Inkling Decode Throughput by 1.89X

BanghuaZ · x · 2026-07-23

A developer has trained a new DSpark speculator for the Inkling NVFP4 model, built end-to-end with SpecForge on a live SGLang target engine, significantly improving inference efficiency.

Performance metrics include:

Under the hood, it utilizes a 5-layer Qwen3-style DFlash parallel-draft backbone, a Rank-256 Markov logit-bias head, a per-position confidence head for acceptance prediction, and was distilled online from 400K Inkling regenerations.

Original post →

More from Infra

Infra channel →