QSA Attention Benchmarked: Better Efficiency and LM Performance
nrehiew_ · x · 2026-08-27
Ablations on short, long, and general LM evals show QSA performs better on language modeling (assuming FLOPs matched) and is on par with RULER. Efficiency-wise, it is significantly better and does not negatively impact MTP acceptance.
Related event: QSA Attention and GatedResidual Outperform Baselines(2 posts)→
More from Research
- Hugging Face incident debate: Model strategy awareness — akbirkhan · 2026-08-27
- Pre-ChatGPT hospital triage chatbot for COVID-19 — AryHHAry · 2026-08-27
- JIT-Agent: Improving LLMs via Just-in-Time Harness Evolution — NationalUniversityofSingapore · 2026-08-27
- D³-MOPD: Dynamic Scheduling for Multi-Teacher Distillation — Zechen Sun · 2026-08-27
- Frontier Models Complete Only ~20% of Scientific Workflows — apodex · 2026-08-27
- Agent-G²: Gaussian Guidance for Long-Horizon RL — ZhejiangUniversity · 2026-08-27