Kimi K3 lands with 2.8T parameters, 1M context, and B200 throughput data
TheZachMueller · x · 2026-08-04
Kimi K3 benchmarked on Lambda with 2× NVIDIA B200s
The post introduces Moonshot AI’s Kimi K3, described as a flagship open-weight model in the 3-trillion-parameter class. It says the model has roughly 2.8T total parameters, about 104B activated per token, a native vision encoder, and a 1-million-token context window.
It also highlights architectural changes over Kimi K2:
- Kimi Delta Attention (KDA), a linear-attention layer with recurrent gated memory
- Gated Multi-head Latent Attention (MLA) to keep million-token context affordable
- Stable LatentMoE with 896 routed experts
On Lambda’s deployment benchmark, the post reports results on 2× NVIDIA HGX B200 with Infiniband and CUDA 13.0:
- 1,491.65 tok/s total throughput
- 298.33 tok/s generation throughput
- 9.32 tok/s per user
- 9,425.53 ms TTFT
- 102.61 ms ITL
The workload used 8,192 input / 2,048 output tokens, 512 prompts, and 32 concurrent requests, aiming to simulate long-context coding and document-analysis usage.
Related event: Moonshot AI Releases Open-Weight Flagship Model Kimi K3(2 posts)→
More from Infra
- Reply says token prices have been falling since late May despite strong demand — GaryMarcus · 2026-08-04
- MiniMax H3 adds open weights, stereo audio and 15-second 2K video generation — petrusenko_max · 2026-08-04
- How to size a local RAG stack for 50 users, from OCR to reranking — InternationalGap3698 · 2026-08-04
- UK sovereign AI fund backs OLIX Computing to rebuild AI infrastructure — HZoete · 2026-08-04
- BlackRock closes $12.5B bond deal for Meta-backed data center — Beth_Kindig · 2026-08-04
- Enterprise AI budgets are already stale as 2027 planning opens — sanjaykalra · 2026-08-04