First systematic study of hybrid linear attention: PAS and ISP activation spikes explained
jiqizhixin · x · 2026-09-10
StartLux, Tsinghua, UCAS, HKU, University of Sydney and Columbia present the first systematic study of Hybrid Linear Attention (HLA) models: 5 linear attention architectures, 6 hybrid configurations, 5 data domains, 12 public checkpoints from 1.2B to 397B parameters (Qwen3.5, Kimi Linear, Nemotron-H, Zamba).
Using Massive Activations as a probe, the team identifies two core phenomena:
- Pre-Attention Spikes (PAS): activations surge to a peak in the layer right before a full attention layer;
- Inter-Spike Plateaus (ISP): sustained high-activation regions between spikes.
The work shows how a few full attention layers reshape computation inside hybrid models.
More from Research
- Hidenori Tanaka's Swarm Interpretability: Why AI Agents Converge on Shared Beliefs Without Rewards — Hidenori8Tanaka · 2026-09-10
- Open-source YuE2 music model generates editable symbolic scores before rendering full songs — GreyScope · 2026-09-10
- Turing launches CEO Bench: 500+ expert tasks test frontier agents in a simulated company — ecekamar · 2026-09-10
- Academics warn AI lets colleagues turn half-baked ideas into papers, breaking incentives further — erikphoel · 2026-09-10
- SpeechLMs secretly transcribe: implicit text-decodable stage found in middle layers — kastnerkyle · 2026-09-10
- 415k hours of full-duplex dialogue speech dataset released for spoken dialogue model training — kastnerkyle · 2026-09-10