Kimi Delta Attention cuts KV cache by 75% and speeds million-token decoding by 6×

johnseach · x · 2026-07-27

How Kimi Delta Attention works

Moonshot/Kimi’s KDA (Kimi Delta Attention) is presented as the key to making Kimi K3 practical at million-token context lengths.

The claimed result is up to 75% less KV cache and up to 6× faster decoding at million-token context, while matching or beating full attention quality in Moonshot’s internal tests.

Original post →

More from Models

Models channel →