LiFT loops a DiT at inference: beats DiT-XL/2 with 52% less inference compute

cgmsnoek · x · 2026-10-06

LiFT (Loop Flow Transformers) turns part of a flow-matching DiT into a recurrent core, letting inference depth scale far beyond training depth with no retraining or early exits. On ImageNet it beats dense DiT-XL/2 by 3.34 FID while using 60% fewer parameters, 32% fewer training FLOPs, and 52% fewer inference FLOPs — a new way to spend inference compute inside each sampling step.

Original post →

More from Multimodal

Multimodal channel →