Debating Whether Positional Encodings Are Truly Needed for Causal Models at Scale

mgostIH · x · 2026-08-09

The discussion centers on the mechanics of linear attention (like DeltaNet) and positional encodings (like RoPE). The quoted post suggests RoPE can be viewed as a product of accumulating transition matrices, which in linear attention can generalize into data-dependent matrices that change at each position.

The replier counters this, arguing that the entire theoretical wall of text might be a delusion: perhaps positional encodings are simply unnecessary for causal models at scale, regardless of the specific architecture used.

Related event: Mathematical Analysis of Kimi K3: Why RoPE Is Dropped(2 posts)→

Original post →

More from Research

Research channel →