Reddit poster argues DeepSeek's Engram is only half of the next attention-scale breakthrough
SrijSriv211 · reddit · 2026-09-10
A Reddit user lays out a personal bet on the next architecture-level breakthrough after attention. Recapping how attention (2017) enabled in-context learning and scaling, the author argues DeepSeek's much-hyped Engram mechanism makes smaller models more knowledgeable but not necessarily smarter — calling it only the first half of the next big breakthrough.
The author's contrarian prediction: the real leap will be architectures where most parameters are input-dependent — small fixed weight matrices generating "fast-weights" on the fly to produce FFN weights dynamically, analogous to how attention itself acts as a fast-weight layer. This, they argue, could yield smaller, smarter models and unlock capabilities currently reserved for much larger networks.
More from AGI Musings
- Anthropic Researcher Puts Over 10% Odds on AI Killing All Humans, BBC Reports — Dan_Jeffries1 · 2026-09-10
- Ex-OpenAI researcher lists 3 copiums on AI progress, still puts simulation theory at 50/50 — AaronBergman18 · 2026-09-10
- Ex-quant trader pivots to AI safety, arguing field is bottlenecked by talent, not money — austinc3301 · 2026-09-10
- AlphaGenome suggests AI's biggest science wins may come from narrowing search spaces — VraserX · 2026-09-10
- sjgadler clashes with Boris: OpenAI's chief scientist sees AI risk differently — sjgadler · 2026-09-10
- Reporter: Anthropic Is a 'Bubble Inside a Bubble' Rarely Discussing Practical AI — sebkrier · 2026-09-10