Open-sourced Kimi K3 speculator lifts single-stream throughput from 118 to 370 tok/s
vllm_project · x · 2026-07-27
- vLLM reports that for a 2.8T model, speculative decoding is the best path to ultra-low latency without accuracy loss.
- Inferact trained and open-sourced a DSpark speculator for Kimi K3 that drafts multiple tokens in one parallel pass.
- Reported throughput improved from 118 tok/s to 370 tok/s on a single stream, a 3.14× gain on real reasoning workloads.
Related event: vLLM brings day-0 support to Moonshot’s Kimi K3(11 posts)→
More from Infra
- fmgo: call Apple's on-device Foundation Models from Go with no CGO and no Swift — Super_Run_8466 · 2026-09-23
- Huawei unveils Peerium architecture: nested BSP unifies million processors into one computer — Dr_Singularity · 2026-09-23
- Grok explains why DeepSeek picked DualPipe + ZeRO-1 over ZeRO-3 on 2048 H800s — TheZachMueller · 2026-09-23
- AI costs fall 47% per quarter, 4x faster than DNA sequencing: Epoch AI — daveholtz · 2026-09-23
- M5 Ultra LLM test: 4x faster prompt processing, but double the power draw — DigitalguyCH · 2026-09-23
- $500 of Dell OptiPlexes become a diskless netboot lab where AI agents can't brick the hardware — colinmcnamara · 2026-09-23