Red Hat Releases New DSpark Models, Boosting vLLM Inference Speed by 4x
vllm_project · x · 2026-08-08
Red Hat AI has released three new DSpark speculators on Hugging Face to accelerate LLM inference through parallel drafting. The models are designed for Qwen3.6-35B-A3B, gemma-4-31B, and GLM-5.2.
DSpark drafts a whole block of tokens in a single pass and then verifies them losslessly. On the GLM-5.2 model, this technique achieves a 4.0x speedup in decoding, boosting single-request generation from 137 to 551 tokens/sec. It also delivers a 1.57x higher peak throughput (from 1,402 to 2,203 tokens/sec) with an acceptance length of 5.47 on math reasoning tasks. These models can be served using @vllmproject.
More from Infra
- Nvidia B300 Specs Contradiction: More SMs but Same BF16 TFLOPS as B200 — StasBekman · 2026-08-08
- Model Routing Reshapes AI Economics: Glean Cuts Latency 50% and Speeds Search 10x — VibeMarketer_ · 2026-08-08
- Running Qwen 3.6 27B on RTX 5090: 40 t/s at 262k Context in llama.cpp — Gargle-Loaf-Spunk · 2026-08-08
- Together AI Releases Interactive Diagrams Explaining LLM Inference and Quantization — zainhas · 2026-08-08
- Nscale Claims $51B Contracted Revenue Ahead of IPO, Faces Industry Skepticism — nathanbenaich · 2026-08-08
- vLLM and NVIDIA Achieve Over 25K TPS/GPU for Qwen3.5 on GB200 Systems — AccBalanced · 2026-08-08