Training-free Recurrent Transformer: Why does top-down activation injection work?
cephaloform · x · 2026-08-19
An interesting architectural finding demonstrates injecting activations from the top layers into the bottom layers at the next time step (Recurrent Transformer). Experiments show that this approach works effectively without training, sparking discussion about the underlying mechanism.
More from Research
- Ornith-1.5 Released: 397B MoE Model Matches Claude Opus Performance — aftahi_ai · 2026-08-20
- Physics of Agents: Studying Opinion Dynamics of 10,000 LLMs — CatAstro_Piyush · 2026-08-20
- Finding a Decades-Old Bug in Knuth's Algorithm D and LLVM — jedisct1 · 2026-08-20
- Red-Blue Pebble Game: Analyzing Matmul Communication Overhead — srush_nlp · 2026-08-20
- Summarizing 1,000 AI Papers Costs Just $4 with DeepSeek V4 Flash — nutlope · 2026-08-20
- Paper: non-LLM components dominate latency in 5 of 10 production agents — dair_ai · 2026-08-20