Kimi K3 launches with 2.8T parameters, 1M context and $0.30 input pricing
togethercompute · x · 2026-07-28
Together’s Kimi K3 launch page adds more concrete product details around the model’s agent use case.
It describes Kimi K3 as Moonshot AI’s most capable model and the first open model in the 3-trillion-parameter class. The page highlights:
- 2.8T total parameters
- 16 of 896 experts activated per token
- Kimi Delta Attention and Attention Residuals as the architectural changes behind the model
- native vision plus a 1-million-token context window
- use cases such as extended engineering sessions, multi-step research, document workflows, and screenshots/charts/documents in the same model
- availability on serverless and dedicated infrastructure with 99.9% SLA
The page also lists launch pricing:
- Cached input: $3.00 per 1M tokens
- Input: $0.30 per 1M tokens
- Output: $15.00 per 1M tokens
Related event: Moonshot's Kimi K3 Flagship Model Launches on Together AI(12 posts)→
More from coding & agent
- Model Is the Least Interesting Part: A Guide to Six Core AI Architectures from RAG to Multi-Agent — goyalshaliniuk · 2026-09-11
- Non-coder builds layered memory architecture: 20k tokens tracks a year of agent conversations — matteoianni · 2026-09-11
- Warp's six non-engineering teams all run on Linear and Claude Code — mon__lim · 2026-09-11
- 9-year backend dev: AI code isn't the problem, the rate of making a mess is — Sweaty-Landscape-561 · 2026-09-11
- RTK claims token savings, but our cost benchmarks disagree — michalwarda · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11