Scaling Kimi K3 on H200s: Engineering Insights from 1000+ Chips
hsu_byron · x · 2026-08-01
Author shares practical tuning insights for high-throughput inference of Kimi K3 on H200 clusters:
- TP+EP beats DP+EP: With only 1/4 of layers using gated MLA, DP attention provides less benefit and increases latency.
- FP8 KV cache works well: No quality or throughput regression measured, effectively doubling cache capacity.
- TP32+EP32 boosts capacity: Provides 4M KV-cache tokens per replica vs 800K with TP16+EP16.
- Disable radix cache for low hit-rate workloads: Saves Mamba-state slots.
- Non-speculative decoding wins at high concurrency: Overhead outweighs savings above 64 concurrent requests.
The setup has run reliably at full load across 1000 chips for days, though some multi-node instability issues are still under investigation.
More from Infra
- Taalas Bakes Llama 3.1 into Custom Silicon, Hitting 15,000 Tokens/sec — generativist · 2026-08-01
- Benchmarks: Running DeepSeek Locally on 4x 5060 Ti with 128k Context — Ambitious_Fold_2874 · 2026-08-01
- 8-Hour Debug of ComfyUI Black Images Uncovers PyTorch FP16 Overflow Bug — rtsitola · 2026-08-01
- Korea's July NAND Export Prices Drop While SSD Prices Rise, Indicating High-Margin Shift — zephyr_z9 · 2026-08-01
- SGLang Supports Inkling-Small on Dual DGX Spark, Hits 24 tok/s — ying11231 · 2026-08-01
- macmon: Open-Source Terminal Performance Monitor for Apple Silicon Hits 1.8k Stars — tom_doerr · 2026-08-01