QSA Attention Benchmarked: Better Efficiency and LM Performance

nrehiew_ · x · 2026-08-27

Ablations on short, long, and general LM evals show QSA performs better on language modeling (assuming FLOPs matched) and is on par with RULER. Efficiency-wise, it is significantly better and does not negatively impact MTP acceptance.

Related event: QSA Attention and GatedResidual Outperform Baselines(2 posts)→

Original post →

More from Research

Research channel →