Kimi K3 launches on SGLang with 423 tok/s and 11 cloud partners
ying11231 · x · 2026-07-28
- Kimi K3 is now live on SGLang with a reported 423 tok/s day-0 speed on GSM8K.
- The post highlights that the demo video was generated by Kimi K3 itself, running on the same serving stack being announced.
- Engineering details include a 2.8T MoE model, a novel KDA architecture, and deep serving optimizations such as fused decode kernels, DP attention, DSpark, PD disaggregation, and KDA-aware prefix caching.
- The ecosystem launch included 11 cloud partners serving the model from day one, framed as a sign of a mature open-source stack where model, infra, and cloud providers ship together.
More from Infra
- Moonshot opensources FlashKDA, claiming 1.72×–2.22× H20 prefill gains — ctjlewis · 2026-07-27
- Claude chat indexing incident sparks a push for confidential-compute AI — bittingthembits · 2026-07-27
- Moonshot open-sources MoonEP as open models vs closed labs debate intensifies — KyeGomezB · 2026-07-27
- Kimi K3 goes live on Nebius with 1M-token context and a 57 AA score — teortaxesTex · 2026-07-27
- AI agent finds a longstanding Bun Node-compat bug in `child_process.spawn` — steipete · 2026-07-27
- llama.cpp adds support for Nanbeige4.2 in pull request 25994 — pmttyji · 2026-07-27