4x Faster Kimi-K3 Decoding: Red Hat Releases DSpark Speculator
woosuk_k · x · 2026-08-14
Red Hat AI has released DSpark, a speculative decoding model optimized for Kimi-K3, delivering massive inference speedups.
- Single-stream interactivity: Boosts speed from 110 to 435 tok/s/user on math reasoning.
- Throughput: Delivers 3.5x the output throughput at matched interactivity under load.
- Long-context stability: Using a 2048-token sliding window attention across all 5 draft layers, the model maintains steady acceptance rates up to 20K context across 10 LongBench domains.
- Built on the vLLM project, it is particularly performant for long-context agentic workloads.
Related event: Red Hat Unveils DSpark Decoder, Greatly Boosting Kimi-K3 Inference(2 posts)→
More from Infra
- Databricks Introduces Smart Routing in Unity AI Gateway, Claims 30%+ Cost Reduction — matei_zaharia · 2026-08-14
- OpenAI Acquired 4.2% Stake in Cerebras Before Ultrafast Launch — ryanmerket · 2026-08-14
- Jensen Huang Marks DGX 10th Anniversary, Unveils DGX Spark at 5x Original Power — nvidia · 2026-08-14
- NVIDIA x Runway: Gen-4.5 Integrated into Vera Rubin Platform in One Day — nvidia · 2026-08-14
- RTX 5090 Test: SageAttention Nearly Doubles Video Generation Speed — gabxav · 2026-08-14
- CoreWeave Sandbox Lets You Drive Claude Agents From Your Phone — wandb · 2026-08-14