Kimi K3 adds native MXFP4 quantization and shows heavy VRAM demands
teortaxesTex · x · 2026-07-27
- The image excerpt says Kimi K3 uses native MXFP4 quantization: MXFP4 weights with MXFP8 activations.
- It claims quantization-aware training starts from the SFT stage to improve hardware compatibility.
- The deployment section says Kimi K3 is accessible through the Kimi platform API and offers OpenAI/Anthropic-compatible APIs.
- It recommends running inference on vLLM, SGLang, or TokenSpeed.
- The accompanying hardware slide says the model’s VRAM needs are large because its MoE experts must be loaded in memory, and gives rough totals such as 480 GB minimum VRAM and 8× H100 as a recommended configuration in one setup.
Related event: Moonshot releases Kimi K3 open weights amid license debate(155 posts)→
More from Infra
- OpenRouter agents now out-consume humans as AI usage arrives in three waves — AccBalanced · 2026-09-11
- Nvidia Is Now Core to Every Major Robotaxi Stack at Commercial Scale — pdamodaran · 2026-09-11
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- RunningHub open-sources H3Lightning, speeding up MiniMax H3 video generation 12x — 智东西 · 2026-09-11