Training-free Recurrent Transformer: Why does top-down activation injection work?

cephaloform · x · 2026-08-19

An interesting architectural finding demonstrates injecting activations from the top layers into the bottom layers at the next time step (Recurrent Transformer). Experiments show that this approach works effectively without training, sparking discussion about the underlying mechanism.

Original post →

More from Research

Research channel →