New attention design called inference-informed: cheap prefill, small KV cache

stochasticchasm · x · 2026-09-11

An analyst assesses a newly revealed attention architecture as 'very inference-informed': cheap prefill, small KV cache, likely friendly for PD disaggregation, with good utilization under larger-batch prefill.

In a follow-up he notes the design resembles hysparse, NSA, and DeepSeek's own CSA/HCA from v4 — combining a local sliding-window branch with a sparse retrieval branch is solidifying as a broad industry pattern.

Related event: Inference-First Architecture Sparks Debate: FP4 KV Cache and Pure CSA2 Compression in Focus(11 posts)→

Original post →

More from Infra

Infra channel →