Open-sourced Kimi K3 speculator lifts single-stream throughput from 118 to 370 tok/s
vllm_project · x · 2026-07-27
- vLLM reports that for a 2.8T model, speculative decoding is the best path to ultra-low latency without accuracy loss.
- Inferact trained and open-sourced a DSpark speculator for Kimi K3 that drafts multiple tokens in one parallel pass.
- Reported throughput improved from 118 tok/s to 370 tok/s on a single stream, a 3.14× gain on real reasoning workloads.
More from Infra
- Claude chat indexing incident sparks a push for confidential-compute AI — bittingthembits · 2026-07-27
- Moonshot open-sources MoonEP as open models vs closed labs debate intensifies — KyeGomezB · 2026-07-27
- Kimi K3’s 2.5x scaling-law gain draws praise for training efficiency — andrew_n_carr · 2026-07-27
- Kimi K3 goes live on Nebius with 1M-token context and a 57 AA score — teortaxesTex · 2026-07-27
- AI agent finds a longstanding Bun Node-compat bug in `child_process.spawn` — steipete · 2026-07-27
- llama.cpp adds support for Nanbeige4.2 in pull request 25994 — pmttyji · 2026-07-27