BAAI's MALA attention allocates its own compute, cutting 128K training latency 2.2x

BAAI · hf · 2026-09-29

BAAI introduces MALA (MassAlloc Attention), a fused attention primitive that keeps score access to every legal causal interaction but allocates post-score computation by normalized contribution. Forward uses an evolving online-softmax normalizer; backward reuses the finalized normalizer; one tolerance governs both training and inference.

Key results:

Original post →

More from Infra

Infra channel →