Databricks Tops Kimi K3 Inference Speed at 239 tokens/s

altryne · x · 2026-08-04

Databricks announced it has secured the #1 spot for Kimi K3 inference speed and latency on Artificial Analysis, achieving 239 tokens/s.

Kimi K3 is a massive 2.8T parameter model, making it the largest open-source model Databricks has ever served. The team highlighted their optimized GPU utilization to achieve this performance.

Related event: Databricks Tops Kimi K3 Inference at 239 tokens/s(2 posts)→

Original post →

More from Infra

Infra channel →