Kimi K3’s 1.4 TB weights fit only on 8×B300, not A100 or H200
qubridInc · reddit · 2026-07-27
Moonshot’s Kimi K3 weights are dropping, and one team is already doing the GPU math before the download lands.
The post says the model is expected to be about 2.8T total parameters, with 896 experts, 16 active experts per token, 1M context, and vision support. The weights are expected to be around 1.4 TB after MXFP4 quantization-aware training.
The author then works through deployment feasibility on three GPU stacks:
- 8× A100 80GB: only 640 GB, so the model does not fit in a single node; dequantization or unsupported INT4 kernels would be required.
- 8× H200: about 1.13 TB, still not enough for a single node.
- 8× B300: about 2.3 TB, enough to fit the full model on one node with room for long-context KV cache.
The model card is also described as unusually candid: quality drops if an agent harness truncates its thinking history, it tends to act instead of asking when ambiguous, and the chat experience is still said to lag models like Fable 5 and Sol even where benchmarks are close.
The team plans to publish throughput, time-to-first-token, and cost-per-million-token results across the three GPU configurations.
Related event: Moonshot releases Kimi K3 open weights amid license debate(155 posts)→
More from Infra
- fmgo: call Apple's on-device Foundation Models from Go with no CGO and no Swift — Super_Run_8466 · 2026-09-23
- Huawei unveils Peerium architecture: nested BSP unifies million processors into one computer — Dr_Singularity · 2026-09-23
- Grok explains why DeepSeek picked DualPipe + ZeRO-1 over ZeRO-3 on 2048 H800s — TheZachMueller · 2026-09-23
- AI costs fall 47% per quarter, 4x faster than DNA sequencing: Epoch AI — daveholtz · 2026-09-23
- M5 Ultra LLM test: 4x faster prompt processing, but double the power draw — DigitalguyCH · 2026-09-23
- $500 of Dell OptiPlexes become a diskless netboot lab where AI agents can't brick the hardware — colinmcnamara · 2026-09-23