BAAI's MALA attention allocates its own compute, cutting 128K training latency 2.2x
BAAI · hf · 2026-09-29
BAAI introduces MALA (MassAlloc Attention), a fused attention primitive that keeps score access to every legal causal interaction but allocates post-score computation by normalized contribution. Forward uses an evolving online-softmax normalizer; backward reuses the finalized normalizer; one tolerance governs both training and inference.
Key results:
- At matched 8K work, MALA approaches a per-instance oracle with mean omitted mass of 0.0188% vs 0.0182%.
- In a 128K-token benchmark with tensor parallelism, training forward/backward latency drops 2.2x and 3.0x, and decoding latency drops 1.6x vs FullAttn.
- Scaling-law runs from 0.6B to 14B track FullAttn perplexity while cutting total training FLOPs; 14B and continued-trained 32B models match FullAttn on knowledge, reasoning, and long-context retrieval.
More from Infra
- SemiAnalysis: Why GLM-5.3 Sparse Attention Doesn't Cut HBM Memory Capacity Needs — burny_tech · 2026-09-29
- Exploit Summit Montreal recap: Gamma tokens, iota SDK, $12M run rate for Targon — markjeffrey · 2026-09-29
- Bain says AI must earn $6T a year by 2031 — matching all global IT spending today — sanjaykalra · 2026-09-29
- On DGX Spark, bf16 beats int8 convrot: H3 video gen 272s vs 287s in real tests — dtdisapointingresult · 2026-09-29
- BAAI's CoWA attention cuts training latency 7.4x while matching FullAttn quality to 32B — BAAI · 2026-09-29
- Databricks Tops All 4 NVIDIA SOL-ExecBench Kernel Tracks Using AI Agents for ~$70K — Yuchenj_UW · 2026-09-29