Expressive power isn't what gradient descent finds: why RNN cells excel at state-tracking generalization

mike64_t · x · 2026-09-14

Argues that traditional recurrent cells have an unreasonably good inductive bias for synthetic state-tracking tasks — they can often literally express the required computation and learn it trivially via teacher forcing, while transformers must build bespoke approximations. Key point: failure to length-generalize on toy tasks doesn't prove an architecture lacks expressive power; what's expressible and what gradient descent finds are different things, especially in sharp loss landscapes.

Original post →

More from Research

Research channel →