Chris Manning notes SMT as joint optimization of memory-bottlenecked Transformers and RNNs

chrmanning · x · 2026-08-25

A substantive machine learning discussion occurred on Twitter. Phillip Isola clarified that while State Space Models (SMT) are sometimes viewed as 'distilled transformers,' they are actually the joint optimization of a memory-bottlenecked transformer and an RNN. This perspective may be closer to the truth than previous categorizations, with details illustrated in Figure 10 of the related paper.

Original post →

More from Research

Research channel →