Only 15 Fused Kernels in Prefill: V4.1's Inference Stack and Minute-Scale SWA Cache

nrehiew_ · x · 2026-09-11

nrehiew sums up inference insights: heavily fused kernels (15 during prefill, 11 during decode), SWA-based KV cache reduction where SWA is cached only at message ends with a minute-scale TTL on DRAM, and a replay scheme bounding max replay by parameter S when caches are evicted. On post-training, he notes the team's philosophy is now completely data-focused rather than training-algorithm research.

Related event: DeepSeek V4.1 Tech Report Deep Dive: RL Infrastructure, Sandbox Design and Inference Stack(8 posts)→

Original post →

More from Infra

Infra channel →