Where Does Memory Live? RNNs, Transformers, and SSMs Compared Through Working Memory

Pretty_Upstairs9035 · reddit · 2026-10-07

The author re-examines three architectures through the lens of working memory: RNNs compress history into a recurrent state (O(N²) parameters carrying only O(N) state), Transformers store the past as KV-cache entries attended over at inference—powerful but frozen weights mean context management rather than durable knowledge—and selective SSMs like Mamba return to fixed-size recurrent memory with input-dependent retention. It highlights BDH (Dragon Hatchling), which pairs linear attention in a high-dimensional neuron space with low-rank implementation, using an N×D recurrent state with Hebbian-like synaptic updates. The author's point: instead of architecture horse races, 'where does memory actually live?' may be the better question—though none of this kills Transformers or solves continual learning.

Original post →

More from Research

Research channel →