LiFT loops a shared DiT to beat dense DiT-XL/2 with 60% fewer parameters on ImageNet
retr0jirachi · x · 2026-10-07
Researchers at the University of Amsterdam introduce Loop Flow Transformers (LiFT), a family of looped generative models that scales inference computation by repeatedly applying a shared Diffusion Transformer core with minimal architectural changes.
Key idea: instead of asking each recurrent step to predict the final target, LiFT trains each step with a single regression target — a point on a straight path from the model's initial estimate to the flow-matching target, indexed by a continuous depth coordinate. A model trained with just 2 loops can loop up to 16 times at inference without retraining, early exits, or other modifications.
Results: on ImageNet 256x256, LiFT-L/2 achieves an FID 3.34 points lower than the dense DiT-XL/2 baseline while using 60% fewer parameters, 32% fewer training FLOPs, and 52% fewer inference FLOPs.
Related event: LiFT: Recurrent DiT Cuts Parameters 60% While Improving FID(3 posts)→
More from Multimodal
- BVB benchmark tests agentic video understanding by rebuilding real videos in Blender — RexDouglass · 2026-10-07
- New tool turns AI-generated 3D models into fully rigged VTuber avatars with VRM export — Promptmethus · 2026-10-07
- Restoration LoRA for LTX-2.5 turns 240p mush into crisp 720p video — multimodalart · 2026-10-07
- Neural Emission Fields: Real-Time Rendering of Pre-integrated Neural Emitters — ssh4net · 2026-10-07
- Blender's The City Generator brings Houdini-style procedural 3D city building to everyone — bilawalsidhu · 2026-10-07
- MiniCPM-V-4.7-35B-A3B quietly appears on Hugging Face without a model card — BreakfastFriendly728 · 2026-10-07