Expressive power isn't what gradient descent finds: why RNN cells excel at state-tracking generalization
mike64_t · x · 2026-09-14
Argues that traditional recurrent cells have an unreasonably good inductive bias for synthetic state-tracking tasks — they can often literally express the required computation and learn it trivially via teacher forcing, while transformers must build bespoke approximations. Key point: failure to length-generalize on toy tasks doesn't prove an architecture lacks expressive power; what's expressible and what gradient descent finds are different things, especially in sharp loss landscapes.
More from Research
- Johns Hopkins preprint finds 3,022 recursive splice sites in the human genome — anshulkundaje · 2026-09-14
- Shanghai team open-sources NCP-ArchPreview, an 8.9B model that predicts concepts instead of tokens — wonderwomancode · 2026-09-14
- Muon-trained models leave a distinctive spectral signature in their weights — QuanquanGu · 2026-09-14
- whitetree: dynamic exact kNN without rebuilds, 40-300x faster than sklearn BallTree — monononon34 · 2026-09-14
- Google AI x Econ team finds field evidence that prior expertise drives learning from AI-assisted work — soumitrashukla9 · 2026-09-14
- Graph theory from hometown streets: navigation, genome assembly and Hamiltonian paths — TivadarDanka · 2026-09-14