Kimi K3 brings million-token context with open-source delta attention

theomitsa · x · 2026-07-26

A quote-thread highlights Kimi K3 as an open model at frontier scale and says it uses a new mechanism called delta attention that avoids a growing KV cache.

The key claim is that this design lets the model handle one million tokens of context without memory usage blowing up. The thread then starts explaining attention from first principles to show why the approach matters.

Related event: Moonshot AI Unveils Trillion-Parameter Kimi K3(4 posts)→

Original post →

More from Models

Models channel →