LIFT: new architecture teaches Transformers deep-to-shallow feedback while keeping training parallel
megamor2 · x · 2026-10-01
New work: Latent Information Feedback Transformers (LIFT). The core problem: in LM generation, information flows from high to low layers only via decoded tokens, creating a bottleneck; removing it via state propagation makes the model recurrent and hard to train in parallel.
- LIFT combines a Transformer-based architecture with teacher-supervised training to exploit deep-to-shallow feedback while keeping pretraining fully parallel.
- Under token-matched budgets, LIFT consistently outperforms standard Transformers and baselines on language modeling, downstream reasoning, and procedural tasks; on par or ahead under compute matching.
- On a state-tracking task, LIFT beats Transformers trained on 8x more data.
More from Research
- Cohere Labs' World Model from Scratch session 3 covers post-training and fine-tuning — Cohere_Labs · 2026-10-02
- Meta's MemLife: training-free egocentric video memory system gains 4.6-12.0%, plus RL-optimized writer MemOpt — meta · 2026-10-02
- Suffix cache reuse for hybrid attention lifts edited-turn hit rate 29.9% to 58.1%, cutting prefix-reuse FLOPs to 7.14/Q — RulinShao · 2026-10-02
- Overmind benchmarks show task-specific SLMs beat frontier models, 7x on contract clause quoting — rohanpaul_ai · 2026-10-02
- Sandia Labs taps Radical's AI self-driving lab to discover new energy materials — CatAstro_Piyush · 2026-10-02
- TimelineBench: Best of 16 AI Agents Passes Just 26.8% of 56 Real Video-Editing Tasks — ycombinator · 2026-10-02