LIFT: teacher-supervised deep-to-shallow feedback for parallel-trained Transformers

megamor2 · x · 2026-10-01

New work introduces Latent Information Feedback Transformers (LIFT): in standard LMs, deep-layer information reaches lower layers only through decoded tokens, a bottleneck; state propagation makes models recurrent and unscalable to train. LIFT uses teacher supervision to give Transformer LMs deep-to-shallow feedback while keeping pretraining fully parallel. Under token-matched budgets, LIFT consistently beats standard Transformers and baselines on language modeling, downstream reasoning, and procedural tasks; it matches or beats compute-matched Transformers, and beats Transformers trained on 8x more data on a state-tracking task.

Related event: LIFT Enables Deep-to-Shallow Feedback in Transformers While Keeping Parallel Training(2 posts)→

Original post →

More from Research

Research channel →