Kimi K3’s 2.8T open weights need data-center hardware to run locally
ArtificialAnlys · x · 2026-07-29
Running Kimi K3 locally needs data-center scale hardware
Artificial Analysis argues that Kimi K3 is far harder to host locally than recent open-weight flagships like GLM-5.2. It says Kimi K3 is the first open-weights model that does not fit on a single Hopper node even at 4-bit precision.
Key points
- 2.8 trillion total parameters make Kimi K3 the largest open-weights model yet.
- At native MXFP4, the weights alone take about 1,560 GB of memory before serving any tokens.
- Its hybrid attention reduces KV-cache demand, so a GB300 NVL72 could theoretically handle 2,000+ concurrent requests at 256K context, versus about 1,000 for Kimi K2.6, if throughput limits are ignored.
- Serving one user at 256K context still needs about 1,570 GB of memory, with roughly 8 GB more per extra concurrent user when using 16-bit KV cache.
- The total system cost would be in the hundreds of thousands of dollars, depending on hardware and pricing.
Practical hardware targets
The post lists realistic configurations that can fit the model and leave room for cache: 8× B300 or GB300, 16× B200 or GB200, or AMD 8× MI355X / MI455X.
Artificial Analysis says it will publish day-one inference benchmarks soon and track how serving performance improves as inference stacks mature.
Related event: Kimi K3 Open Weights Demand Data Center Hardware(3 posts)→
More from Infra
- Qualcomm goes agent-centric: Snapdragon 8 Elite Gen 6 and agent-native devices — jiqizhixin · 2026-09-23
- Unsloth Desktop Hotfix Adds Qwen-Image-2.1 Image Editing and Fixes GGUF Loading — danielhanchen · 2026-09-23
- Qwen 3.6 35B-A3B Q6 hits ~50 tok/s on a 128GB Strix Halo — what's the best local model now? — jankeydankey · 2026-09-23
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23