How Kimi K3 engineered its way to the frontier: Delta Attention, expert balancing, AgentENV
noninertialframe96 · reddit · 2026-07-31
Moonshot's Kimi K3 reached frontier level, ranked 4th of 580 models by Artificial Analysis (behind Claude Opus 5, Fable 5, GPT-5.6 Sol). Technical report highlights:
- Kimi Delta Attention: replaces KV cache in 69/93 layers with a 128x128 matrix per head; 1M-token context uses 27.2 GiB vs 104.6.
- Quantile Balancing: keeps 896 experts/layer evenly loaded by computing bias directly.
- AgentENV: Firecracker microVM runtime created 51M sandboxes with 133ms checkpoints and 49ms resumes.
Full walkthrough: https://codepointer.substack.com/p/how-kimi-k3-engineered-its-way-to
More from Models
- OpenAI Slashes GPT-5.6 Small Model Prices by Up to 80% Amid Business Cost Scrutiny — Polymarket · 2026-07-31
- OpenAI Slashes GPT-5.6 Series Prices, Adds Fast Mode for Sol Model — daniel_mac8 · 2026-07-31
- Kimi K3 Successfully Runs Locally on a Consumer Basement PC — Yuchenj_UW · 2026-07-31
- GPT-5.6 Luna Price Slashed by 80% as OpenAI Sparks API Price War — sama · 2026-07-31
- Google DeepMind Launches Gemini Robotics 2: Single Model Controls Whole-Body Movement of Humanoids — DynamicWebPaige · 2026-07-31
- OpenAI Confirms Leaked HF Model Isn't GPT-6, Tested Lowering Cyber Refusals — rickasaurus · 2026-07-31