BAAI's CoWA attention cuts training latency 7.4x while matching FullAttn quality to 32B

BAAI · hf · 2026-09-29

FullAttn repeatedly exposes the complete causal history to every attention head, wasting compute and memory traffic even with IO-efficient dense kernels. BAAI's CoWA distributes access to causal history across KV heads: all heads share near-diagonal and prefix-sink windows, while complementary long-range windows partition the remaining history. The union gives full causal coverage without learned routers or indexers, stays consistent between training and inference, and aligns with KV-head tensor parallelism.

A window-matched ablation at 8K shows CoWA with 100% collective coverage reaches 89.73% accuracy vs 89.97% for FullAttn, while duplicated long-range windows fare much worse; in matched-budget associative recall, CoWA tracks FullAttn as context grows where other sparse patterns degrade.

On a 128K attention-operator benchmark, CoWA cuts training forward/backward latency 7.4x/8.6x and decoding latency 3.0x, with 7.6x lower per-rank decoding memory. Scaling-law runs from 0.6B to 14B track FullAttn perplexity at lower total FLOPs, and a separately continued-trained 32B model matches FullAttn on knowledge, reasoning, and long-context retrieval.

Original post →

More from Infra

Infra channel →