Looped-DiT: 260M looped model beats 6.5x larger text-to-image rival with 4.9x less compute

Apprehensive_Sky892 · reddit · 2026-10-04

Looped Diffusion Transformer (Looped-DiT) scales text-to-image models by re-running shared Transformer blocks within each denoising step, increasing compute depth without adding parameters.

Key findings:

The authors argue looped computation is a more effective scaling axis for visual generation than raw model size.

Related event: Looped-DiT beats 6.5x larger models with 260M parameters via layer reuse(3 posts)→

Original post →

More from Multimodal

Multimodal channel →