LIFT: teacher-supervised deep-to-shallow feedback for parallel-trained Transformers
megamor2 · x · 2026-10-01
New work introduces Latent Information Feedback Transformers (LIFT): in standard LMs, deep-layer information reaches lower layers only through decoded tokens, a bottleneck; state propagation makes models recurrent and unscalable to train. LIFT uses teacher supervision to give Transformer LMs deep-to-shallow feedback while keeping pretraining fully parallel. Under token-matched budgets, LIFT consistently beats standard Transformers and baselines on language modeling, downstream reasoning, and procedural tasks; it matches or beats compute-matched Transformers, and beats Transformers trained on 8x more data on a state-tracking task.
More from Research
- A visual refresher on the basics of Markov chains — alexbilz · 2026-10-01
- A $1 million prize for scientific honesty could reshape research culture — skdh · 2026-10-01
- Agent0: zero-data self-evolving agent framework from Stanford/Salesforce headed to COLM2026 — yuyinzhou_cs · 2026-10-01
- Melting Pot updated: Lab2d ships modern Python wheel, no more sandboxed old versions — jzl86 · 2026-10-01
- NormViz benchmark: best model Gemini 3 Flash scores just 25.3% on visual cultural norms across 16 countries — StellaLisy · 2026-10-01
- Workspace Models: Memory Architecture Turns Reasoning Agents into Robot Policies — jajoosam · 2026-10-01