Kimi K3 Hits Together AI: 3T Parameter MoE with 1M Token Context
zainhas · x · 2026-07-28
Moonshot AI's latest model, Kimi K3, is now available on the Poe platform via Together AI.
Based on platform details, Kimi K3 introduces the following key features:
- Scale & MoE: Boasts 2.8 trillion total parameters (3T class) with a Mixture-of-Experts (MoE) architecture, activating 16 out of 896 experts per token.
- Architectural Changes: Incorporates Kimi Delta Attention and Attention Residuals, modifying information flow across sequence length and model depth.
- Native Capabilities: Features native vision (capable of reading screenshots, charts, and documents) and a massive 1-million-token context window.
- Use Case: Designed for long-horizon tasks such as extended engineering sessions across large repositories and multi-step research. It launches with maximum thinking effort enabled and open-licensed weights.
Related event: Moonshot's Kimi K3 Flagship Model Launches on Together AI(12 posts)→
More from Models
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- Fully local voice assistant on an RTX 3060 replicates the GPT Live demo in 6.5 minutes — liampetti · 2026-09-11
- Meta's Muse Agent has built-in invite code logic, hinting at free-usage expansion — testingcatalog · 2026-09-11