Kimi K3 report reveals training and systems stack
围绕月之暗面 Kimi K3 的多条帖子,作者们集中拆解了其技术报告,信息覆盖训练流程、MoE 架构、KDA 推理与并行、多模态视觉塔以及内部评测。按这些帖子转述,K3 不只给出了编码、Agent、对话和 WebDev 方向的内测表,还公开了 FlashKDA、MoonEP、双缓存推理、快照式沙箱等系统层设计。对行业的看点在于,这是一份把模型效果、训练方法和底层工程一起摊开的技术披露。
已确认
- 多条帖子称,K3 的训练流程为 SFT→RL→MOPD;其中 SFT 使用 QAT,权重为 MXFP4、激活为 MXFP8。
- 架构层面,帖子转述 K3 采用了 LatentMoE、用于专家负载均衡的“分位数均衡”、按每 12 层分组的块注意力残差,以及 NoPE 机制。
- 系统实现上,帖子称 K3 使用 FlashKDA 让 KDA 支持 chunkwise parallel;为控制数值稳定性,还会进一步切成 16 个 token 的 tile。推理侧同时处理 KDA state 与 MLA KV cache,并配有快照式沙箱。
- 多模态方面,帖子称其视觉塔 MoonViT-V2 并非先用 SigLIP 一类模型初始化,而是与语言模型一起从头训练,训练目标仍是 next-token prediction。
- 评测与案例方面,分享材料提到的内部表格覆盖 coding、agent、dialogue 与 WebDev;案例研究则包括 GPU kernel 优化、编译器和芯片原型。@nrehiew 还转述称,较早期的 K3 checkpoint 在开发后期已能承担团队大部分 kernel 优化工作。
尚未确认
- @nrehiew 根据 FLOPs 曲线推测,K3 的 RL 计算量在专家之间可能并不均匀,且大量算力也许花在 MOPD 上;这是其个人分析,不是帖子里能直接坐实的官方结论。
- @nrehiew 还根据流水线图判断 K3 使用标准 1F1B 而非 Dual Pipe,且 EP all-to-all 做了常规重叠处理;这同样属于读图解读。
为什么重要
- 这些帖子之所以引发关注,不只是因为 K3 被拿来与多款闭源、开源模型同台比较,更因为其技术披露覆盖了从量化训练、MoE 均衡到推理缓存和多模态联合训练的完整链路。
- @nrehiew 与另一位开发者都把这份报告评价为少见地“把细节讲清楚”的材料。对关注超大规模 MoE、原生多模态和高效推理基础设施的读者来说,这批信息的参考价值高于单纯跑分。
2026-07-28 ~ 2026-07-29 · 19 related posts
- Kimi K3 case studies show kernel optimizations, a Triton-like compiler, and a chip prototype — teortaxesTex · 2026-07-28
- Kimi K3 Optimization: KDA Numerical Stability and Chunkwise Parallelism — nrehiew_ · 2026-07-29
- Kimi K3 Architecture: Block Attention Residuals and NoPE Integration — nrehiew_ · 2026-07-29
- Kimi K3 Architecture Analysis: Quantile Balancing and LatentMoE for 3T Scale — nrehiew_ · 2026-07-29
- Kimi K3 trains MoonViT-V2 from scratch to stabilize multimodal training — nrehiew_ · 2026-07-29
- Kimi K3 trains its vision encoder from scratch and claims better vision evals — nrehiew_ · 2026-07-29
- Kimi K3 claims 2.5× scaling efficiency gains from KDA savings — nrehiew_ · 2026-07-29
- Kimi K3 synthesizes RL tasks from a web-search-built knowledge graph — nrehiew_ · 2026-07-29
- Kimi K3 uses QAT, RL and MOPD across a wide expert-task mix — nrehiew_ · 2026-07-29
- Kimi K3 Infrastructure: FlashKDA and Chunkwise Parallel Optimization — nrehiew_ · 2026-07-29
- Kimi K3 Parallelism Strategy: Integrating ViT into the Pipeline Diagram — nrehiew_ · 2026-07-29
- Kimi K3 adds MoonEP balancing and ViT pipeline parallelism to its training stack — nrehiew_ · 2026-07-29
- Kimi K3 uses FlashKDA to make chunkwise parallel context computation work — nrehiew_ · 2026-07-29
- Kimi K3 report details dual cache inference and snapshot-based sandbox infra — nrehiew_ · 2026-07-29
- Kimi K3 report adds in-house coding, agent, and WebDev benchmark tables — nrehiew_ · 2026-07-29
- Kimi K3 reportedly helped with kernel optimization and beats several models on in-house benches — nrehiew_ · 2026-07-29
- Dev Reviews Kimi K3: Solid Tech Report, Handles Kernel Optimization — nrehiew_ · 2026-07-29