BAAI's CoWA attention cuts training latency 7.4x while matching FullAttn quality to 32B
BAAI · hf · 2026-09-29
FullAttn repeatedly exposes the complete causal history to every attention head, wasting compute and memory traffic even with IO-efficient dense kernels. BAAI's CoWA distributes access to causal history across KV heads: all heads share near-diagonal and prefix-sink windows, while complementary long-range windows partition the remaining history. The union gives full causal coverage without learned routers or indexers, stays consistent between training and inference, and aligns with KV-head tensor parallelism.
A window-matched ablation at 8K shows CoWA with 100% collective coverage reaches 89.73% accuracy vs 89.97% for FullAttn, while duplicated long-range windows fare much worse; in matched-budget associative recall, CoWA tracks FullAttn as context grows where other sparse patterns degrade.
On a 128K attention-operator benchmark, CoWA cuts training forward/backward latency 7.4x/8.6x and decoding latency 3.0x, with 7.6x lower per-rank decoding memory. Scaling-law runs from 0.6B to 14B track FullAttn perplexity at lower total FLOPs, and a separately continued-trained 32B model matches FullAttn on knowledge, reasoning, and long-context retrieval.
More from Infra
- SemiAnalysis: Why GLM-5.3 Sparse Attention Doesn't Cut HBM Memory Capacity Needs — burny_tech · 2026-09-29
- Exploit Summit Montreal recap: Gamma tokens, iota SDK, $12M run rate for Targon — markjeffrey · 2026-09-29
- Bain says AI must earn $6T a year by 2031 — matching all global IT spending today — sanjaykalra · 2026-09-29
- On DGX Spark, bf16 beats int8 convrot: H3 video gen 272s vs 287s in real tests — dtdisapointingresult · 2026-09-29
- BAAI's MALA attention allocates its own compute, cutting 128K training latency 2.2x — BAAI · 2026-09-29
- Databricks Tops All 4 NVIDIA SOL-ExecBench Kernel Tracks Using AI Agents for ~$70K — Yuchenj_UW · 2026-09-29