Kimi K3 report details dual cache inference and snapshot-based sandbox infra
nrehiew_ · x · 2026-07-29
The thread describes Kimi K3’s efficient inference stack, focusing on how it handles two cache types at once: KDA state and MLA KV cache.
- Because KDA uses a recurrent state shape, block-level hashing does not work well, so the system uses different hashing granularities for MLA and KDA.
- A figure shows fine-grained prefix caching within a 6144-token physical block, where persisted KDA checkpoints are sparse and tied to conversation boundaries.
- The poster also notes K3’s sandbox infrastructure, highlighting secure execution and efficient pause/resume from snapshots.
- The post further mentions offloading external KV cache to CPU as prefix cache and training states to NVMe, though some details are uncertain in the thread.
More from Infra
- DGX Spark arrives as a local AI workstation for LimestoneHQ experiments — alex_verem · 2026-07-29
- Atlassian caps employee AI spend at up to $2,000 a month — nordicinst · 2026-07-29
- The Evolution of Enterprise AI Retrieval: Indexing Becomes Key to Accuracy and Latency — damianplayer · 2026-07-29
- Europractice MPW 2025 chart shows TSMC leads foundry designs with 251 — jwt0625 · 2026-07-29
- A practical Krea 2 LoRA guide targets 1024-res training on 16GB VRAM rigs — Endlesswoodtrail · 2026-07-29
- Kernel Forge uses MCTS to optimize CUDA kernels and beats PyTorch baselines on 14 cases — omarsar0 · 2026-07-29