Kimi K3 Parameters and Inference Margins Revealed
zephyr_z9 · x · 2026-07-19
The post mentions that Kimi K3 has a total of about **2.8T** parameters, which scales down to roughly **1.4T** when served in **fp4**. It is called one of the sparsest frontier models on the market, with a sparsity of only **1.7%**. It also provides an inference-side estimate: its inference business has a gross margin of at least **75%–85%**. The following quote adds a perspective on DeepSeek—even if model weights and parts of the architecture are open-sourced, competitors still struggle to replicate its real cost advantages because key implementation details and operational experience remain proprietary.
Related event: Moonshot Releases 2.8T Open-Weight Model Kimi K3(13 posts)→
More from Infra
- TokenPrint turns Qwen inference into a DevTools-style visual debugger — Rich-Fruit-326 · 2026-07-21
- Remote AI model costs may push companies toward owning their own LLMs — DavidLinthicum · 2026-07-21
- A broken agent router burned 30.2M tokens in 3.5 hours on Claude Code — RileyRalmuto · 2026-07-21
- Huawei's Atlas 950 SuperPoD Scales to 500,000 Chips with Unified Architecture — pstAsiatech · 2026-07-21
- South Korea's exports jump 50% in early July on the AI chip boom — Polymarket · 2026-07-21
- Bittensor boosters argue decentralized training can offset severalfold compute gaps — markjeffrey · 2026-07-21