Kimi K3 Available on Baseten with vLLM-Powered Production API
vllm_project · x · 2026-07-30
The vLLM project officially announced that the Kimi K3 model is now available via Baseten's production API. Baseten uses vLLM as part of its inference stack, allowing users to call the 2.8-trillion-parameter model without managing the underlying infrastructure themselves.
More from Infra
- QuixiAI Open-Sources SlimServe: Fast Inference for GLM on AMD MI300X — QuixiAI · 2026-07-30
- ThunderAgent by Together AI: 2x Faster Agentic Inference (ICML 2026) — togethercompute · 2026-07-30
- Together Compute unveils multi-node inference engine with near-linear scaling, adopted by SkyRL and NVIDIA Dynamo — togethercompute · 2026-07-30
- DigitalOcean Becomes Day 0 Inference Partner for Kimi K3 with 1M Context — vllm_project · 2026-07-30
- Kimi K3 Launches with Day 0 Support on AMD Instinct via vLLM — vllm_project · 2026-07-30
- Debunking the DeepSeek and Chinese Lithography Panic: Exaggerated Costs and Gaps — teortaxesTex · 2026-07-30