Looped-DiT: 260M looped model beats 6.5x larger ones with 4.9x less inference compute
burny_tech · x · 2026-10-04
A new paper from SenseTime, Tsinghua, NTU and collaborators proposes Looped Diffusion Transformer (Looped-DiT), an alternative scaling path for diffusion models: repeatedly running shared Transformer blocks within each denoising step to increase computational depth without growing the parameter count.
- Naive looping fails to consistently improve image quality due to weak supervision across intermediate loops and unregulated attention updates that erode local information
- The fix combines deep supervision across loops with self-modulating attention to stabilize feature updates
- Under matched-parameter and matched-compute settings, a 260M looped model surpasses a model 6.5x larger across text-to-image benchmarks while using 4.9x less inference compute
- Under a fixed inference budget, increasing loop depth yields larger gains than adding denoising steps; deeper loops progressively correct earlier mistakes, suggesting latent reasoning
The authors argue looped computation is a promising way to scale visual generation models.
Related event: Looped-DiT: Layer Reuse Beats Models 6.5x Larger with 4.9x Less Compute(2 posts)→
More from Multimodal
- AI Creator Riffs Lord of the Rings as Rapper Gandalf x DJ Frodo with Higgsfield — TinfoilTricorn · 2026-10-04
- 10 wild examples of Opus 5.5 building games, 3D worlds, videos and ads — minchoi · 2026-10-04
- Adding characters to scenes fails in Qwen: broken scale and degraded faces — rcplaybox · 2026-10-04
- Brainstorm-to-Casting-Room workflow turns AI characters into reusable assets — dr_cintas · 2026-10-04
- Fable 5.5 generates a stunning 40,000-year art history animation — dotey · 2026-10-04
- Augie's AI Image Browser indexes ComfyUI metadata and offers leak-free exports for selling — spanktastic0x · 2026-10-04