Moonshot’s Kimi K3 is a 2.8T open-weight MoE model with 1M-token context

alex_verem · x · 2026-07-29

Moonshot’s Kimi K3 technical report argues that the next leap in AI may come from using compute more efficiently, not just scaling it up.

What the model is

Why the team says it matters

The report claims roughly 2.5× better intelligence per unit of compute versus Kimi K2.

Reported results

Post-training reinforcement learning across general, coding, and reasoning tasks is said to improve compositional generalization and long-horizon execution.

At 2.8T scale, the system is presented as a combined effort in algorithm design, KDA-balanced expert-parallel training, memory management, million-token agentic RL, persistent rollout, sandbox states, and deployment innovations.

Benchmarks and positioning

Original post →

More from Infra

Infra channel →