Flash-MSA: Open-Sourced Sparse Attention Accelerates Million-Token Training

The open-source sparse attention kernel Flash-MSA has been released to optimize long-context training efficiency. By replacing dense attention with a customized kernel, it specifically accelerates training for sequences at the million-token scale.

2026-07-13 ~ 2026-07-13 · 2 related posts