Databricks tops Kimi K3 (max) benchmark with 218.8 t/s and 10.31s latency
mrdrozdov · x · 2026-08-04
Key findings
Artificial Analysis benchmarked 10 API providers for Kimi K3 (max) across output speed, time to first token, and blended price.
- Databricks ranked first on both output speed at 218.8 tokens/s and lowest latency at 10.31s.
- Wafer (FAST) came second on both metrics, with 194.5 tokens/s and 11.14s.
- Fireworks followed at 166.7 tokens/s and 13.14s latency.
- On price, the lowest blended cost was $2.31 / 1M tokens, shared by Kimi, Fireworks, Modal, Together AI, and DigitalOcean.
The report says provider choice materially changes performance, with a 637% spread between the fastest and slowest providers.
More from Infra
- Ibiden’s AI substrate pricing surge sets up a clean earnings asymmetry — tengyanAI · 2026-08-04
- Big Tech’s OpenAI and Anthropic stakes are inflating reported earnings — Kr00ney · 2026-08-04
- Menlo Ventures says AI has entered phase 2, with infrastructure as the real opportunity — mmurph · 2026-08-04
- Podcast says AI CapEx, compute crunch, and debt-financed data centers are squeezing semis — BenBajarin · 2026-08-04
- MiniMax H3 open weights run 32 minutes down to 7.2 minutes on an L40S — ashishsanu · 2026-08-04
- Fluidstack takes its AI infrastructure dinner series to Austin and keeps hiring — MxMnr · 2026-08-04