Reddit poster argues DeepSeek's Engram is only half of the next attention-scale breakthrough

SrijSriv211 · reddit · 2026-09-10

A Reddit user lays out a personal bet on the next architecture-level breakthrough after attention. Recapping how attention (2017) enabled in-context learning and scaling, the author argues DeepSeek's much-hyped Engram mechanism makes smaller models more knowledgeable but not necessarily smarter — calling it only the first half of the next big breakthrough.

The author's contrarian prediction: the real leap will be architectures where most parameters are input-dependent — small fixed weight matrices generating "fast-weights" on the fly to produce FFN weights dynamically, analogous to how attention itself acts as a fast-weight layer. This, they argue, could yield smaller, smarter models and unlock capabilities currently reserved for much larger networks.

Original post →

More from AGI Musings

AGI Musings channel →