LiFT loops a DiT at inference: beats DiT-XL/2 with 52% less inference compute
cgmsnoek · x · 2026-10-06
LiFT (Loop Flow Transformers) turns part of a flow-matching DiT into a recurrent core, letting inference depth scale far beyond training depth with no retraining or early exits. On ImageNet it beats dense DiT-XL/2 by 3.34 FID while using 60% fewer parameters, 32% fewer training FLOPs, and 52% fewer inference FLOPs — a new way to spend inference compute inside each sampling step.
More from Multimodal
- Kandinsky 6: MIT open-weight video+audio model with 6 variants, day-0 Diffusers support — RisingSayak · 2026-10-06
- One Photo to Talking-Head Reel: Dev Shares $0.50/Video AI Pipeline — victor_explore · 2026-10-06
- AnimeGen ships: run the Anima model locally and offline on iPhone/iPad — Agitated-Pea3251 · 2026-10-06
- Reusable 'Quantum Kaleidoscope City' prompt template for shifting-city image generation — LudovicCreator · 2026-10-06
- One-shotting a new wiki interface with Opus, illustrated by FLUX.1 on Cloudflare — Kippi3000 · 2026-10-06
- Sarvam AI demos end-to-end video dubbing that keeps original voices with line-level edits — itsOmSarraf_ · 2026-10-06