NVIDIA’s Dynamo docs add Kimi-K3 deployment recipes for GB200 and GB300

woosuk_k · x · 2026-07-29

NVIDIA’s Dynamo docs now include a deployment guide for Moonshot AI’s Kimi-K3.

The page describes serving Kimi-K3 on GB200 or GB300 with Dynamo + vLLM, using either aggregated or prefill/decode-disaggregated topologies. It highlights a multimodal MoE model with up to 1M-token context, KV-aware routing, MXFP4-packed routed experts, FP8 KV cache, and NVLink-based transfer and tensor parallelism on NVL72 racks.

Original post →

More from Infra

Infra channel →