Moonshot Releases Open-Weight Kimi K3 Model
Moonshot AI officially released Kimi K3 along with its open weights and a comprehensive technical report. The model boasts a total parameter count of 2.8T, utilizing a Mixture-of-Experts (MoE) architecture that activates 1040 亿参数 per token. It supports a 1 million token context window and features native visual understanding. While the release adds a major player to the ultra-large open-weight model camp, its specific commercial licensing terms have sparked intense debates within the developer community.
已确认
- 模型规模与架构: Kimi K3 has a total of 2.8T parameters with 896 experts, activating 16 (approximately 1040 亿参数). Architecturally, it introduces Hybrid Attention combining Kimi Delta Attention (KDA) and Gated MLA, and replaces the traditional SwiGLU gating with SiTU-GLU to enhance training stability. Additionally, KDA incorporates a lower-bounded decay design to resolve overflow issues under finite precision.
- 性能与能力: Officially claimed as the world's first open 3T-class model, it possesses native multimodal and Agent capabilities. Regarding security tasks, the technical report notes that the model discovered 16 previously unknown vulnerabilities (including 2 Linux kernel vulnerabilities) and outperformed GLM-5.2 in related exploitation tasks. It also scored 91 on the BrowseComp benchmark.
- 部署门槛: The model weights are massive at 1.4TB. Self-hosting requires at least 8 B300 GPUs (with some estimates suggesting 18 GPUs) to fully accommodate it.
尚未确认
- 商业化许可争议: Kimi K3 adopts a Modified MIT license. The terms stipulate that if a user or its affiliates generate a total revenue exceeding 2000 万美元 in any consecutive 12-month period and resell it as a cloud service, they must sign a separate commercial agreement with Moonshot AI. Users like @Eyelbee criticized this for forcing inference providers to raise prices close to official closed-source API levels, severely undermining the accessibility of open-weight models.
为什么重要
Kimi K3 showcases cutting-edge engineering exploration in architectural design (such as three-axis scaling and KDA improvements) and multi-teacher post-training strategies. Its practical vulnerability mining capabilities also validate the application potential of ultra-large models in the security domain. However, Moonshot's high commercial threshold for "Model as a Service" establishes a highly controversial new paradigm for the commercialization of open-source models, which will directly impact the willingness of third-party cloud inference vendors to participate.
2026-07-26 ~ 2026-07-28 · 147 related posts
Primary sources
- Moonshot AI Launches Open-Weight Kimi K3 to Rival Top US Models — emmanuelvivier · 2026-07-26
- Moonshot AI launches open-weight Kimi K3 and claims strong results against top US models — emmanuelvivier · 2026-07-26
- Moonshot’s Kimi K3 ships open weights, but self-hosting needs 1.4TB and 18+ GPUs — Common_Dream9420 · 2026-07-27
- [source] Kimi K3’s 1.4 TB weights fit only on 8×B300, not A100 or H200 — qubridInc · 2026-07-27
- Moonshot AI publishes Kimi-K3 on Hugging Face — mervenoyann · 2026-07-27
- Kimi-K3 is now live on Hugging Face — himanshustwts · 2026-07-27
- Kimi K3 weights are now available — SavunOski · 2026-07-27
- Moonshot makes Kimi K3 officially open weights with a 1M-token context window — scaling01 · 2026-07-27
- Moonshot releases Kimi K3, a 2.8T multimodal MoE with 1M-token context — 月之暗面 Kimi · 2026-07-27
- Moonshot's Kimi K3 repository goes live with a 2.8T-parameter model card — tokenbender · 2026-07-27
- Kimi K3 adds native MXFP4 quantization and shows heavy VRAM demands — teortaxesTex · 2026-07-27
- [source] Moonshot's Kimi K3 license allows broad use, but adds commercial thresholds above $20M — natolambert · 2026-07-27
- Kimi K3 License Breakdown: Non-Commercial Limits for Firms Over $20M Revenue — natolambert · 2026-07-27
- Kimi K3 hardware estimates point to 480 GB VRAM and 8× H100 setups — cedric_chee · 2026-07-27
- Moonshot publishes Kimi-K3 on Hugging Face with 2.8T parameters and 1M context — BankApprehensive7612 · 2026-07-27
- Hugging Face model card details Kimi K3’s 2.8T-parameter multimodal setup — op7418 · 2026-07-27
- Moonshot opens Kimi K3 weights and report, citing a 2.8T MoE model — eliebakouch · 2026-07-27
- Moonshot launches Kimi K3 with 2.8T parameters, 1M context, and native multimodality — mervenoyann · 2026-07-27
- Kimi K3 license sets $20M/year inference terms and a separate $20M/month product deal — natolambert · 2026-07-27
- Kimi K3 weights are now available on Hugging Face — inductionheads · 2026-07-27
- Kimi K3 is now available on Hugging Face — Kyrannio · 2026-07-27
- Moonshot says Kimi K3 is now out — _akhaliq · 2026-07-27
- Moonshot’s Kimi K3 report says its largest model has 2.78T parameters and 2.5× better scaling — teortaxesTex · 2026-07-27
- Kimi K3 License Analysis: Large Enterprises Require Specific Commercial Deals — himanshustwts · 2026-07-27
- Moonshot’s Kimi-K3 model page is now live on Hugging Face — iamfakhrealam · 2026-07-27
- Moonshot updates Kimi K3 license but withholds day-one support for workers.ai — michellechen · 2026-07-27
- Cloudflare adds Moonshot’s Kimi K3 to AI Gateway, but Workers AI lacks day-one support — michellechen · 2026-07-27
- Kimi K3 triples parameters and doubles active experts in a new MoE design — teortaxesTex · 2026-07-27
- Kimi K3’s 2.5x scaling-law gain draws praise for training efficiency — andrew_n_carr · 2026-07-27
- Moonshot releases Kimi K3 report, claiming 2.5× scaling-efficiency gain over K2 — cedric_chee · 2026-07-27
- Kimi K3 reportedly improves training efficiency by 2.5× — zephyr_z9 · 2026-07-27
- Kimi K3 moves from MIT to a new license tied to MaaS revenue thresholds — AdinaYakup · 2026-07-28
- Kimi-K3’s license is free for self-hosting, but clouds must share profits — scaling01 · 2026-07-28
- Kimi K3 report says numerical stability remains a hard problem at scale — teortaxesTex · 2026-07-28
- Attention Residuals may expose cross-layer information flow directly for interpretability — tokenbender · 2026-07-28
- Kimi K3 claims 2.5× better scaling efficiency with a three-axis architecture — suchenzang · 2026-07-28
- K3 Architecture Enables Direct Observation of Cross-Block Routing for Model Interpretability — tokenbender · 2026-07-28
- Moonshot’s Kimi K3 lands on Hugging Face as a 2.8T MoE open model — gaganghotra_ · 2026-07-28
- Kimi K3 arrives as a 2.8T MoE with 104B active params and a 1M-token context — stochasticchasm · 2026-07-28
- Kimi K3 report omits hardware details, leaving its 2.5× efficiency claim hard to verify — cedric_chee · 2026-07-28
- Moonshot Releases Kimi K3 Tech Report, Cites 2.5x Scaling Efficiency Gain — cedric_chee · 2026-07-28
9 near-duplicate retellings: teortaxesTex · HarveenChadha · _akhaliq · iamfakhrealam · ricklamers · cyb3rops · ns123abc · FlorianGallwitz · andrew_n_carr