Kimi K3 launches as a 2.8T open-weights model with 1M context and $3 input pricing
The Batch (Andrew Ng) · rss · 2026-07-22
Main headline: Kimi K3
Moonshot released Kimi K3, a 2.8 trillion-parameter open-weights model with native vision, a 1 million-token context window, and a sparse MoE design that activates 16 of 896 experts.
- The company says two architecture changes, Kimi Delta Attention and Attention Residuals, improve scaling efficiency by about 2.5x versus K2.
- API pricing is live now at $3 / million input tokens and $15 / million output tokens; cached coding input can drop to $0.30 / million tokens.
- Full weights are scheduled for release by July 27, 2026.
- Moonshot demoed K3 on autonomous GPU compiler development, chip design in 48 hours, and automated video editing, while acknowledging it still trails Claude Fable 5 and GPT 5.6 Sol in head-to-head tests.
Other items in the same issue
- Thinking Machines released Inkling, a 975B-parameter open-weights MoE model with 41B active parameters, trained on 45T tokens across text, images, audio, and video.
- Inkling emphasizes controllable thinking effort and matches Nvidia’s Nemotron 3 Ultra on agentic coding tasks with about one-third of the tokens.
- The EU ordered Google to open Android more to rival AI services and share anonymized search data, after finding non-Google agents could not function at Gemini’s level on Android.
- Nvidia launched Nemotron 3 Embed, a family of open embedding models; the 8B model tops RTEB at 78.5%, and the 1B variants are optimized for production and Blackwell GPUs.
- Google renamed NotebookLM to Gemini Notebook.
- Hugging Face said it used GLM 5.2 to help analyze a weekend-long AI agent intrusion that its commercial frontier models refused to inspect.
More from Companies & People
- a16z partner flips to call for nationalizing frontier AI labs, sparking debate — S_OhEigeartaigh · 2026-09-11
- Glean Is Worth $7.2B, but What's Actually Its Moat? — yogthinks · 2026-09-11
- Replit CEO joins backlash against AI doomers, says focus should be on cybersecurity not sci-fi — amasad · 2026-09-11
- Replit CEO Amjad Masad: AI existential-risk talk is a distraction from real cybersecurity work — amasad · 2026-09-11
- SF hiring of wet-lab biologists for RL environments draws bio-safety concerns — beffjezos · 2026-09-11
- Schmidhuber appears on the "Europe on Edge" podcast — SchmidhuberAI · 2026-09-11