LiFT loops a shared DiT to beat dense DiT-XL/2 with 60% fewer parameters on ImageNet

retr0jirachi · x · 2026-10-07

Researchers at the University of Amsterdam introduce Loop Flow Transformers (LiFT), a family of looped generative models that scales inference computation by repeatedly applying a shared Diffusion Transformer core with minimal architectural changes.

Key idea: instead of asking each recurrent step to predict the final target, LiFT trains each step with a single regression target — a point on a straight path from the model's initial estimate to the flow-matching target, indexed by a continuous depth coordinate. A model trained with just 2 loops can loop up to 16 times at inference without retraining, early exits, or other modifications.

Results: on ImageNet 256x256, LiFT-L/2 achieves an FID 3.34 points lower than the dense DiT-XL/2 baseline while using 60% fewer parameters, 32% fewer training FLOPs, and 52% fewer inference FLOPs.

Related event: LiFT: Recurrent DiT Cuts Parameters 60% While Improving FID(3 posts)→

Original post →

More from Multimodal

Multimodal channel →