Looped DiT: 260M model beats 6.5x larger rival on T2I benchmarks with 4.9x less compute

arankomatsuzaki · x · 2026-10-01

A new paper introduces Looped Diffusion Transformer (Looped-DiT), an alternative scaling path for text-to-image models: instead of growing parameters or denoising steps, it repeatedly runs shared Transformer blocks within each denoising step, deepening computation at fixed parameter count. Naive looping fails due to weak supervision and unregulated attention updates, so the authors add deep supervision across loops and self-modulating attention. A 260M-parameter looped model outperforms a 6.5x larger model across T2I benchmarks while using 4.9x less inference compute. Under fixed budgets, deeper loops yield larger gains than more denoising steps and can progressively correct earlier mistakes.

Original post →

More from Multimodal

Multimodal channel →