Kimi 3 Parameters and Attention Mechanism Revealed

bookwormengr · x · 2026-07-16

The post claims that Kimi 3 has a total of 2.8T parameters.

The author finds the most interesting point to be its use of Kimi Delta attention to avoid massive KV cache overhead, while also featuring native vision capabilities.

He ranks its performance just behind Fable and GPT-5.6 Sol, adding that "it will catch up after post-training."

Related event: Kimi K3 Debuts Strong, Narrowing the Open-Weight Gap(184 posts)→

Original post →

More from Models

Models channel →