Looped DiT: 260M model beats 6.5x larger rival on T2I benchmarks with 4.9x less compute
arankomatsuzaki · x · 2026-10-01
A new paper introduces Looped Diffusion Transformer (Looped-DiT), an alternative scaling path for text-to-image models: instead of growing parameters or denoising steps, it repeatedly runs shared Transformer blocks within each denoising step, deepening computation at fixed parameter count. Naive looping fails due to weak supervision and unregulated attention updates, so the authors add deep supervision across loops and self-modulating attention. A 260M-parameter looped model outperforms a 6.5x larger model across T2I benchmarks while using 4.9x less inference compute. Under fixed budgets, deeper loops yield larger gains than more denoising steps and can progressively correct earlier mistakes.
More from Multimodal
- Aiden Bai shares a prompt that generates a 30-second game trailer with original music and sound effects — aidenybai · 2026-10-01
- Google Arts & Culture launches nom nom: a Neural Cellular Automata world simulated by Gemini — zzznah · 2026-10-01
- Stanford's UniEvo-VL Self-Distillation Lifts Qwen-image GenEval From 0.747 to 0.808 — stanfordnlp · 2026-10-01
- Early user feedback: Suno v6 disappoints on music generation quality — hq4ai · 2026-10-01
- Runway teases mysterious Project Continuum with early preview at AI Summit — c_valenzuelab · 2026-10-01
- Viggle Turbo v0.3 for Qwen Image: cleaner 6-step output, new 9-step mode fixes small text — init-5 · 2026-10-01