Databricks Tops Kimi K3 Inference Speed at 239 tokens/s
altryne · x · 2026-08-04
Databricks announced it has secured the #1 spot for Kimi K3 inference speed and latency on Artificial Analysis, achieving 239 tokens/s.
Kimi K3 is a massive 2.8T parameter model, making it the largest open-source model Databricks has ever served. The team highlighted their optimized GPU utilization to achieve this performance.
Related event: Databricks Tops Kimi K3 Inference at 239 tokens/s(2 posts)→
More from Infra
- XM paper leads one reader to expect compute prices to keep rising — jfischoff · 2026-08-04
- How do you water-cool 8 RTX 6000 Pros and a 500W CPU without leaks? — __JockY__ · 2026-08-04
- Data centers don’t just circulate water—they discharge it — djcows · 2026-08-04
- Gemma 4 26B runs locally on a 13-year-old Xeon CPU with zero GPUs — davidad · 2026-08-04
- AI video creators look for high-VRAM GPU alternatives after RTX 5090 shortages — equanimous11 · 2026-08-04
- ARC competitor says its CUDA C stack is an order of magnitude ahead — jsuarez · 2026-08-04