AWS Guide: Deploying 2.8T Parameter Kimi K3 Requires 8x B300 GPUs

AWS ML Blog · rss · 2026-07-31

AWS published a detailed guide for deploying the Kimi K3 model on its cloud infrastructure. Kimi K3 is a 2.8 trillion parameter Mixture of Experts (MoE) model with 104 billion active parameters per token and a 1 million token context window.

The post outlines two primary deployment paths:

Infrastructure Requirements: Deploying this model requires the ml.p6-b300.48xlarge instance equipped with 8 NVIDIA B300 GPUs. The recommended serving engine is a dedicated vLLM container, and model weights are distributed in MXFP4 format for memory efficiency.

Original post →

More from Infra

Infra channel →