Looped-DiT: 260M looped model beats 6.5x larger ones with 4.9x less inference compute

burny_tech · x · 2026-10-04

A new paper from SenseTime, Tsinghua, NTU and collaborators proposes Looped Diffusion Transformer (Looped-DiT), an alternative scaling path for diffusion models: repeatedly running shared Transformer blocks within each denoising step to increase computational depth without growing the parameter count.

The authors argue looped computation is a promising way to scale visual generation models.

Related event: Looped-DiT: Layer Reuse Beats Models 6.5x Larger with 4.9x Less Compute(2 posts)→

Original post →

More from Multimodal

Multimodal channel →