2-Layer Recurrent Networks Match 32-Layer Feedforward Baselines at Same Compute, Thread Claims
mike64_t · x · 2026-09-26
A research thread argues that with just two layers, a recurrent network that remembers the previous step can match a 32-layer feedforward baseline under identical compute budgets.
The author's takeaway: the field may be spending compute on the wrong axis — scaling depth instead of memory. The full thread expands on the argument and experiments.
More from Research
- ICML 2027 PC Asks for Ideas on Handling AI Slop and Review Overload — MarkSchmidtUBC · 2026-09-26
- Jevless: Jev-style typed decisions from any model with logprobs, plus a local server — seraphius · 2026-09-26
- Physicist reveals terms of his paid Anthropic blog post: mid-market rate, no NDA signed — burny_tech · 2026-09-26
- Paper 'Mecha-nudges for Machines' Accepted as NeurIPS 2026 Spotlight — ethayarajh · 2026-09-26
- AutoScreen: AI Agents Reprioritize CRISPR Hits, Reveal How Cancer Cells Evade Immune Attack — KexinHuang5 · 2026-09-26
- Fixing one prompt failure can silently break another: output budgets make prompt patches zero-sum — ClickOk5811 · 2026-09-26