DSpark speculator trained on live SGLang lifts decode throughput 1.89× on B200s

ying11231 · x · 2026-07-23

A team says it trained a DSpark speculator for Inkling NVFP4 using SpecForge against a live SGLang serving engine, instead of an offline dataset.

What changed

Reported results

Under the hood

The team says the serving command is in the comments for anyone who wants to try it.

Related event: Open-Sourced DSpark Speculator Boosts Inkling Throughput by 1.89x(3 posts)→

Original post →

More from Infra

Infra channel →