NVIDIA’s Dynamo docs add Kimi-K3 deployment recipes for GB200 and GB300
woosuk_k · x · 2026-07-29
NVIDIA’s Dynamo docs now include a deployment guide for Moonshot AI’s Kimi-K3.
The page describes serving Kimi-K3 on GB200 or GB300 with Dynamo + vLLM, using either aggregated or prefill/decode-disaggregated topologies. It highlights a multimodal MoE model with up to 1M-token context, KV-aware routing, MXFP4-packed routed experts, FP8 KV cache, and NVLink-based transfer and tensor parallelism on NVL72 racks.
More from Infra
- SK hynix reportedly signs long-term contracts with about 10 major customers — dejavucoder · 2026-07-29
- Shanghai Aishengna is said to be manufacturing DUV lithography tools — zephyr_z9 · 2026-07-29
- A vendor-agnostic Vulkan backend cuts edge inference latency from 30 ms to 3 ms — ppchaos · 2026-07-29
- pdf-mcp turns technical PDFs into structured text, images, and searchable context — tom_doerr · 2026-07-29
- Moonshot’s Kimi K3 is a 2.8T open-weight MoE model with 1M-token context — alex_verem · 2026-07-29
- New scaling law paper says repetition can beat paraphrasing for some pretraining regimes — burny_tech · 2026-07-29