Kimi K3 hardware estimates point to 480 GB VRAM and 8× H100 setups
cedric_chee · x · 2026-07-27
- The image breaks down the hardware requirements for Kimi K3, describing it as the world’s largest open-source model.
- It estimates a minimum 480 GB VRAM requirement under Q2K quantization, with 8× H100 as a recommended configuration.
- The slide explains the large memory footprint by noting that Kimi K3 is a MoE model with 896 experts, even though only 16 are active per token.
- It gives rough memory math for multiple precisions, including 5,600 GB at FP16, 1,540 GB at INT4, and 980 GB at Q2K, before noting expert parallelism can reduce the per-GPU load.
- The image also points to a no-GPU API alternative via the Moonshot platform.
More from Infra
- Cloudflare adds Moonshot’s Kimi K3 to AI Gateway, but Workers AI lacks day-one support — michellechen · 2026-07-27
- Baseten adds Kimi K3 to its Model APIs on launch day — baseten · 2026-07-27
- MoonshotAI open-sources MoonEP for perfectly balanced expert parallelism — teortaxesTex · 2026-07-27
- Moonshot releases Kimi K3, a 2.8T MoE model with 1M context and 423 tok/s serving — ricklamers · 2026-07-27
- How to run a private local AI on many 16GB laptops with Ollama and Gemma 4 — minchoi · 2026-07-27
- Vercel adds Kimi K3 on US providers with zero-data-retention support — cramforce · 2026-07-27