Recurrent Looped Transformer solves 256-bit parity at 100% where standard Transformers stay at chance
princetonu · hf · 2026-10-08
Princeton's Recurrent Looped Transformer (RLT) splits layers between a parallel causal encoder and a recurrent decoder, so computation per token grows with sequence length at fixed per-token cost. Trained on at most 40 bits, RLT generalizes parity to 256 bits at 100% accuracy across seeds while an 8-layer Transformer stays at chance. On swap-based S5 tracking at 8x training length, RLT hits 97% vs under 1%; on modular arithmetic it reaches 93% vs 33%. Ablations confirm the feedback loop is essential, and chunked feedback keeps parity but breaks permutation tracking.
More from Research
- Wikidata Search Traces: 10,235 traces for training knowledge graph search agents — omarsar0 · 2026-10-08
- NVIDIA study: tool use cuts multimodal model refusals of harmful requests by up to 68.7% — JeremyCMorgan · 2026-10-08
- New OpenAI paper extends Riemann zeta zero-free region to Re(s) > 7/8 — PTenigma · 2026-10-08
- OpenAI math paper on Weil classes reported flawed, raising doubts about unformalized proofs — ctjlewis · 2026-10-08
- ZooWork-ShopRanker: open e-commerce rerankers (0.6B-8B) aligned to shopping preferences — kalyan_kpl · 2026-10-08
- LessWrong Thought Experiment: How Should a Model Guess Today's Date With No Date Context? — LessWrong 精选 · 2026-10-08