Kimi K3’s 1.4 TB weights fit only on 8×B300, not A100 or H200
qubridInc · reddit · 2026-07-27
Moonshot’s Kimi K3 weights are dropping, and one team is already doing the GPU math before the download lands.
The post says the model is expected to be about 2.8T total parameters, with 896 experts, 16 active experts per token, 1M context, and vision support. The weights are expected to be around 1.4 TB after MXFP4 quantization-aware training.
The author then works through deployment feasibility on three GPU stacks:
- 8× A100 80GB: only 640 GB, so the model does not fit in a single node; dequantization or unsupported INT4 kernels would be required.
- 8× H200: about 1.13 TB, still not enough for a single node.
- 8× B300: about 2.3 TB, enough to fit the full model on one node with room for long-context KV cache.
The model card is also described as unusually candid: quality drops if an agent harness truncates its thinking history, it tends to act instead of asking when ambiguous, and the chat experience is still said to lag models like Fable 5 and Sol even where benchmarks are close.
The team plans to publish throughput, time-to-first-token, and cost-per-million-token results across the three GPU configurations.
Related event: Moonshot Releases Open-Weight Kimi K3 Model(155 posts)→
More from Infra
- Podcast spotlights the future of vector databases in the AI stack — CShorten30 · 2026-07-29
- Qdrant, Future AGI and AWS set a talk on self-improving agent retrieval loops — qdrant_engine · 2026-07-29
- Google opens early-access Gemini distillation service for smaller, cheaper models — ccerrato147 · 2026-07-29
- A 27B model reaches 24 TPS with on-the-fly 3-bit dequantization on an A6000 — cephaloform · 2026-07-29
- OpenRouter’s weighted average token price fell sharply this year — maferase · 2026-07-29
- Decentralized Network Trains 16B Model Across 3 Continents Using RTX 4090s — bittingthembits · 2026-07-29