2-Layer Recurrent Networks Match 32-Layer Feedforward Baselines at Same Compute, Thread Claims

mike64_t · x · 2026-09-26

A research thread argues that with just two layers, a recurrent network that remembers the previous step can match a 32-layer feedforward baseline under identical compute budgets.

The author's takeaway: the field may be spending compute on the wrong axis — scaling depth instead of memory. The full thread expands on the argument and experiments.

Original post →

More from Research

Research channel →