LiFT: looped DiT cuts params 60% while improving FID by 3.34 on ImageNet
VISLab-Amsterdam · hf · 2026-10-06
VISLab Amsterdam introduces Loop Flow Transformers (LiFT), a family of looped generative models that scale computation by repeatedly applying a shared Diffusion Transformer core:
- Each recurrent step is trained with a single regression target — a point on a straight path from the model's initial estimate to the flow-matching target — indexed by a continuous depth coordinate, so a trained model can loop far beyond its training depth with no retraining or early exits.
- Longer rollouts keep improving generation, meaning inference compute can grow without adding parameters.
- On ImageNet 256×256, LiFT-L/2 achieves FID 3.34 lower than a dense DiT-XL/2 baseline with 60% fewer parameters, 32% fewer training FLOPs, and 52% fewer inference FLOPs.
Related event: LiFT: Loop Flow Transformers Cut Inference Compute 52% While Improving FID(2 posts)→
More from Multimodal
- Runway announces World Runner: generative worlds on a pocket Game Boy-style console — c_valenzuelab · 2026-10-06
- Gothic vampire short film made with Seedance 2.5 on Runwayml — azed_ai · 2026-10-06
- AI-generated superheroes just want chai and biscuits, not saving the world — umesh_ai · 2026-10-06
- YarnGPT quietly ships Pidgin voice translation, more languages coming — saheedniyi_02 · 2026-10-06
- Doubao and Qwen voice models clone timbre from ~10s of audio, cheap enough for agents — AlchainHust · 2026-10-06
- Kling 4.0 Flash dialogue quality impresses: characters now act with pauses and eye contact — LudovicCreator · 2026-10-06