A small technical thread asks whether attention residuals can still benefit from positional embeddings
stochasticchasm · x · 2026-07-28
- The author suggests that attention residuals may also be “NoPE” and asks whether positional embeddings can still matter for them.
- A reply notes that earlier linear-attention + NoPE hybrids still used some RoPE in global layers, raising the question of whether this is truly the first such hybrid.
- The thread is a small but technical discussion about how positional information interacts with attention residual paths.
Related event: K3 and Attention Residuals Advance LLM Interpretability(3 posts)→
More from Research
- Kimi K3 bounds decay at -5 to keep chunkwise KDA inside BF16 range — suchenzang · 2026-07-28
- Block Attention Residuals cuts attention overhead from O(Ld) to O(Nd) — stochasticchasm · 2026-07-28
- Celltype pitches LLMs that predict biological response and is hiring in New York — david_van_dijk · 2026-07-28
- Claude Opus 5 keeps Opus 4.8 pricing while matching Fable 5 within 0.5% on coding — AlexKim · 2026-07-28
- MLA Architecture Details: Will Full-Rank Gate Projection Cause Parameter Explosion? — stochasticchasm · 2026-07-28
- Kimi K3 replaces KDA’s low-rank output gate with an input-dependent full-rank projection — stochasticchasm · 2026-07-28