Moonshot opens Kimi K3 weights, technical report, and key training infra
智东西 · wechat · 2026-07-28
Moonshot AI has open-sourced Kimi K3’s model weights, technical report, and key training infrastructure, including MoonEP, FlashKDA, and AgentEnv.
K3 is presented as a 2.8T-parameter MoE model with native vision and a 1 million-token context window. Moonshot says the model is now freely downloadable and deployable for both internal research and user-facing products.
What was released
- MoonEP: a high-performance communication library for large sparse MoE systems
- FlashKDA: a high-performance kernel for KimiDeltaAttention, with reported 1.72–2.22× prefill speedup on Nvidia H20 versus a flash-linear-attention baseline
- AgentEnv: a sandbox system for large-scale agent training, built with KVCache.ai
Reported benchmarks and capabilities
- Strong results on long-horizon coding and agent tasks such as SWEMarathon, ProgramBench, TerminalBench2.1, BrowseComp, AutomationBench, and SpreadsheetBench2
- Better performance on several coding and knowledge-work tests, though still behind Claude Fable 5 on some broader office-style benchmarks
- Strong visual understanding on chart and document tasks, including OmniDocBench
Ecosystem reaction
Cognition said K3 is now integrated into Devin desktop and CLI. Nebius, Baseten, Fireworks, and Huawei Ascend CANN also announced day-0 support or deployment adaptation. The release drew unusually fast community attention on Hugging Face and X.
Related event: Moonshot Releases 2.8T Parameter Open-Weight Model Kimi K3(44 posts)→
More from Infra
- SovereignAI says Qwen3.5-397B can rival Opus 4.8 for about $450k in compute — schwarzjn_ · 2026-07-28
- Together AI says Kimi K3 targets long tool-heavy agent workflows with 1M context — togethercompute · 2026-07-28
- Cloud and AI prices may keep rising under quarter-on-quarter growth pressure — DavidLinthicum · 2026-07-28
- Falling Inference Compute Costs Could Make 'Vibe Hacking' Very Cheap — joshua_saxe · 2026-07-28
- Kimi K3 throughput jumps from 19 to 49 tok/s on OpenRouter — cedric_chee · 2026-07-28
- llama.cpp adds Kimi-K3 text model support for local inference — ilintar · 2026-07-28