Databricks says Kimi K3 now runs at 239 tokens per second on its serving stack

Yuchenj_UW · x · 2026-08-04

Databricks says it now serves Kimi K3 at 239 tokens per second

Databricks says it has become the #1 provider for Kimi K3 inference speed and latency on Artificial Analysis, with the model reaching 239 tokens/s.

Original post →

More from Infra

Infra channel →