Red Hat Releases New DSpark Models, Boosting vLLM Inference Speed by 4x

vllm_project · x · 2026-08-08

Red Hat AI has released three new DSpark speculators on Hugging Face to accelerate LLM inference through parallel drafting. The models are designed for Qwen3.6-35B-A3B, gemma-4-31B, and GLM-5.2.

DSpark drafts a whole block of tokens in a single pass and then verifies them losslessly. On the GLM-5.2 model, this technique achieves a 4.0x speedup in decoding, boosting single-request generation from 137 to 551 tokens/sec. It also delivers a 1.57x higher peak throughput (from 1,402 to 2,203 tokens/sec) with an acceptance length of 5.47 on math reasoning tasks. These models can be served using @vllmproject.

Original post →

More from Infra

Infra channel →