Looped-DiT: 260M Looped Model Beats 6.5x Larger Diffusion Models With 4.9x Less Compute
NandoDF · x · 2026-10-03
- New paper Looped Diffusion Transformer (Looped-DiT) scales computation by rerunning shared Transformer blocks within each denoising step, adding depth without adding parameters.
- Naive looping fails; the authors trace it to weak supervision on intermediate loops and unregulated attention updates that erode local information.
- Fixes: deep supervision across loops plus self-modulating attention to stabilize feature updates.
- A 260M-parameter looped model surpasses a 6.5x larger model on text-to-image benchmarks while using 4.9x less inference compute.
- Under fixed inference budgets, increasing loop depth yields larger gains than adding denoising steps, and deeper loops progressively correct earlier mistakes.
Related event: Looped-DiT: Layer Reuse Beats Models 6.5x Larger with 4.9x Less Compute(2 posts)→
More from Research
- EgoTools: 100-hour egocentric tool-use dataset lifts an 8B model's accuracy by 10.9 points — liuziwei7 · 2026-10-04
- ARROW unifies 3D reconstruction and point tracking from arbitrary image sets, sets new SOTA — CSProfKGD · 2026-10-04
- Looped-DiT: 260M looped model beats 6.5x larger ones with 4.9x less inference compute — burny_tech · 2026-10-04
- Latent communication papers: KV-cache integrity vs malicious agents and state-delta hybrid channels — burny_tech · 2026-10-04
- DNA design trick: fit with gradient descent first, then snap to manufacturer catalogue — anshulkundaje · 2026-10-04
- Nature study finds synaptic plasticity in biological neural networks approximates backpropagation — aran_nayebi · 2026-10-04