Baseten Breaks Down Serving 2.8T-param Kimi K3 at Scale on Blackwell
thursdai_pod · x · 2026-08-09
Philip Kiely from Baseten provided an in-depth breakdown of how to serve massive 2.8-trillion parameter models like Kimi K3 at scale using Blackwell GB300s.
To achieve this highly challenging deployment, the team performed deep optimizations on the underlying inference stack and contributed their improvements back to the open-source community, enhancing the capabilities of vLLM and SGLang for ultra-large models.
More from coding & agent
- v0-mcp: Generate React UI Components from Natural Language and Design Images — modelcontextprotocol · 2026-08-09
- Qwen 4B Model Breaks Down in Under 30 Seconds Inside Agent Simulated Environment — JayB_Official · 2026-08-09
- Open-Sourcing Oil Motion: AI Video Skill for Interactive Web Animations — churchkey · 2026-08-09
- Developer Test: Using Claude Opus to Automate Code Documentation — EricBuess · 2026-08-09
- SWE-bench Creator on AI Coding: Complex Tooling Is Becoming Obsolete — jyangballin · 2026-08-09
- Combining herder and tmux for a Seamless LLM Workflow on Linux — Rasmic · 2026-08-09