Study reveals cross-layer activation patterns in hybrid attention models

机器之心 · wechat · 2026-09-01

Researchers from StartLux, Tsinghua University, and others have systematically characterized the "cross-layer traces" left by sparse FullAttention layers in Hybrid Linear Attention LLMs (HLALLM) using a probe called "MassiveActivations."

Key Findings: Spikes and Plateaus

Mechanism: Write—Sink—Cancel

The study proposes a three-stage lifecycle to explain this phenomenon:

Broad Validation

This pattern is verified in public models like KimiLinear, Qwen3.5, and Nemotron-H (1.2B-397B) and emerges early in training. The findings suggest that FullAttention layout acts as an "organizer" for the internal computational rhythm, offering new insights for model compression and quantization.

Original post →

More from Research

Research channel →