Recurrent Looped Transformer solves 256-bit parity at 100% where standard Transformers stay at chance

princetonu · hf · 2026-10-08

Princeton's Recurrent Looped Transformer (RLT) splits layers between a parallel causal encoder and a recurrent decoder, so computation per token grows with sequence length at fixed per-token cost. Trained on at most 40 bits, RLT generalizes parity to 256 bits at 100% accuracy across seeds while an 8-layer Transformer stays at chance. On swap-based S5 tracking at 8x training length, RLT hits 97% vs under 1%; on modular arithmetic it reaches 93% vs 33%. Ablations confirm the feedback loop is essential, and chunked feedback keeps parity but breaks permutation tracking.

Original post →

More from Research

Research channel →