Databricks says Kimi K3 now runs at 239 tokens per second on its serving stack
Yuchenj_UW · x · 2026-08-04
Databricks says it now serves Kimi K3 at 239 tokens per second
Databricks says it has become the #1 provider for Kimi K3 inference speed and latency on Artificial Analysis, with the model reaching 239 tokens/s.
- Kimi K3 is described as a 2.8T-parameter model and the largest open-source model Databricks has served so far.
- The post highlights Databricks’ GPU throughput and positioning around high-performance serving.
- The attached chart shows Databricks leading other providers on latency/speed for Kimi K3.
More from Infra
- llama.cpp patches boost DeepSeek-V4-Flash-0731 from 3.26 to 25.91 tok/s — dyn___ · 2026-08-04
- Inference provider swings Kimi-K3 benchmark results, with one endpoint topping CEO-Bench — AAAzzam · 2026-08-04
- Hyperscalers' AI backlog hits $2.3T, but analyst warns of circular financing — TiernanRayTech · 2026-08-04
- Formula 1 cuts data-source onboarding from 8 weeks to 40 minutes with AWS agents — AWS ML Blog · 2026-08-04
- Anthropic explores India data residency with AWS as Claude eyes in-country inference — HimanshiET · 2026-08-04
- Google Cloud adds borderless Lakehouse for Gemini Enterprise across AWS, Databricks and Snowflake — rseroter · 2026-08-04