Debating Whether Positional Encodings Are Truly Needed for Causal Models at Scale
mgostIH · x · 2026-08-09
The discussion centers on the mechanics of linear attention (like DeltaNet) and positional encodings (like RoPE). The quoted post suggests RoPE can be viewed as a product of accumulating transition matrices, which in linear attention can generalize into data-dependent matrices that change at each position.
The replier counters this, arguing that the entire theoretical wall of text might be a delusion: perhaps positional encodings are simply unnecessary for causal models at scale, regardless of the specific architecture used.
Related event: Mathematical Analysis of Kimi K3: Why RoPE Is Dropped(2 posts)→
More from Research
- ChatGPT finds normalization error in two Riemann Hypothesis papers, author confirms — theimposingshadow · 2026-08-09
- MIT Professor Likens AI Discovery Process to Physics Phase Transitions — CatAstro_Piyush · 2026-08-09
- Eterna Launches Million-Scale RNA Self-Replicating Molecule Design Quest — chaitjo · 2026-08-09
- MiniMax H3 Acceleration v0.2.1: Offline Smoothing Replay Fixes Audio Quality Loss — marres · 2026-08-09
- Researcher Slams AI Eval Orgs for Lacking Basic Software Engineering Practices — evijit · 2026-08-09
- AI solves 25-year-old open problem in wireless communications: polynomial-time algorithm found — DimitrisPapail · 2026-08-09