Kimi K3 adds native MXFP4 quantization and shows heavy VRAM demands
teortaxesTex · x · 2026-07-27
- The image excerpt says Kimi K3 uses native MXFP4 quantization: MXFP4 weights with MXFP8 activations.
- It claims quantization-aware training starts from the SFT stage to improve hardware compatibility.
- The deployment section says Kimi K3 is accessible through the Kimi platform API and offers OpenAI/Anthropic-compatible APIs.
- It recommends running inference on vLLM, SGLang, or TokenSpeed.
- The accompanying hardware slide says the model’s VRAM needs are large because its MoE experts must be loaded in memory, and gives rough totals such as 480 GB minimum VRAM and 8× H100 as a recommended configuration in one setup.
More from Infra
- Cloudflare adds Moonshot’s Kimi K3 to AI Gateway, but Workers AI lacks day-one support — michellechen · 2026-07-27
- Baseten adds Kimi K3 to its Model APIs on launch day — baseten · 2026-07-27
- MoonshotAI open-sources MoonEP for perfectly balanced expert parallelism — teortaxesTex · 2026-07-27
- Moonshot releases Kimi K3, a 2.8T MoE model with 1M context and 423 tok/s serving — ricklamers · 2026-07-27
- How to run a private local AI on many 16GB laptops with Ollama and Gemma 4 — minchoi · 2026-07-27
- Vercel adds Kimi K3 on US providers with zero-data-retention support — cramforce · 2026-07-27