Looped Diffusion Transformer
Yong Xien Chng, Tianyi Chen, Wenwen Tong, Haiwen Diao, Zhongang Cai, Lei Yang, Ziwei Liu, Lewei Lu, Dahua Lin, Gao Huang
cs.CV, cs.LG
2026-10-01
Looped-DiT reuses shared Transformer blocks N times within each denoising step, with deep supervision and self-modulating attention. A 260M text-to-image model tops a 1.7B baseline across six benchmarks at 4.9x lower inference compute.
Text-to-image models have two standard ways to scale: more parameters (Qwen-Image and FLUX.2 sit in the billions) or more denoising steps. Both push up deployment cost. A third option comes from the language-model side: looped Transformers, which apply the same middle blocks to hidden states repeatedly. Depth grows; parameter count and sequence length do not.
For image generation the appeal goes beyond compute. A coherent image requires inferring implied content, resolving interdependent constraints, and catching inconsistencies mid-generation. These are inherently visual processes, and looping performs them directly in hidden representations, with no explicit text reasoning trace.
Naive looping does not work, though. Applied to MiniT2I, a minimal pixel-space MMDiT, it can sit below the non-looped baseline at shallow depths, saturate, then degrade past the training loop count. The authors probe this with a ridge-regression decoder that predicts each image token's 2D grid position from its hidden state: R² falls from 0.865 after loop 1 to 0.562 after loop 8. Repeatedly applying the same transformation lets attention updates overwrite and erode local information in the image tokens.
Two root causes, two fixes: intermediate loops get no direct supervision (gradients must travel through every later loop), and attention update strength never adapts across loops.
The 17 MMDiT blocks are split into a pre-loop stage A, a middle group of 5 blocks B run N times (N=4 in training), and a post-loop stage C. B's parameters are shared across loops, so extra loops add compute depth, not parameters. Looping happens inside each denoising step, decoupled from the sampler's step count.
Two components fix naive looping:
The ablations back the mechanism, not just the scores. XSA keeps the attention-update norm relative to the residual stream smallest and position decodability highest; without it, later loops keep editing an already-correct image and add an extra object. XSA gains +1.6 on the looped model versus −0.1 and +0.7 on parameter- and compute-matched non-looped baselines, so the modulation matters specifically under repeated updates.
Deep supervision has a useful side effect: the model tolerates 1 to 8 inference loops despite training at 4, and even its single-loop setting beats the non-looped baseline. That enables adaptive looping. A lightweight gating network (cross-attention with 4 learnable queries plus an MLP head) predicts the expected benefit of another loop; one threshold λ slides the average loop count from 1.25 to 2.70 without retraining, beating fixed-loop inference at matched average loops.
Average scores across six benchmarks:
| Model | Params | GenEval | DPG | PRISM | CoRe | Spatial | TIIF-Short | Avg |
| MiniT2I-L/16 | 1.6B | 90.3 | 83.4 | 58.5 | 45.0 | 51.6 | 78.2 | 67.8 |
| InternVL-U | 1.7B | 85.0 | 85.2 | 63.5 | 48.2 | 54.5 | 77.7 | 69.0 |
| Uni-CoT (best CoT) | 6.6B | 81.2 | 84.1 | 58.8 | 45.8 | 53.8 | 76.8 | 66.8 |
| Looped-DiT B/16 | 0.26B | 87.4 | 87.0 | 67.0 | 53.5 | 54.6 | 79.7 | 71.5 |
Versus the next-best model, InternVL-U, that is +2.5 average with 6.5x fewer parameters, 4.9x lower inference compute, and 6.5x lower latency. It leads on five of six benchmarks; GenEval goes to MiniT2I-L/16.
The gains survive three control comparisons. On the B/32 variant under matched parameters and compute, looping adds +4.0 over the parameter-matched baseline and reaches 59.1 versus 58.1 for a deeper baseline also given deep supervision (matched training compute). Under a fixed inference budget, spending compute on loop depth beats spending it on more denoising steps consistently across 25-to-50-step settings. Against textual CoT prompt rewriting, looping wins big on constraint-satisfaction subtasks (PRISM/Long Text +9.1), CoT wins on inferring implied content (CoRe/Generalization +22.2), and combining both scores highest.
Score rises monotonically from loop 1 to 4. Qualitatively, successive loops fill in missing content, reorganize objects to meet spatial constraints, and fix rendering errors that persist in the larger non-looped and explicit-CoT baselines shown.
This is the first systematic study of looped computation for text-to-image generation, and the evidence chain is thorough: gains hold under matched parameter, inference-compute, and training-compute settings, and adding depth, width, denoising steps, or CoT does not reproduce them. For practitioners, looping is an inference-time scaling knob orthogonal to model size and step count, and one checkpoint slides along the latency-quality curve by loop count alone.
XSA is a parameter-free change worth trying anywhere computation repeats. Adaptive looping's retrain-free compute control has obvious uses for tiered serving.
The honest caveat: this is a controlled study on MiniT2I at 260M scale. Whether the conclusions transfer to industrial multi-billion-parameter models is untested.
Stated by the authors: naive looping is unreliable and needs intermediate supervision plus adaptive regulation; the reasoning evidence is behavior "suggestive of" latent reasoning, with no mechanistic account of what the loops compute.
Questions from reading the paper. The headline comparison uses MiniT2I-series baselines trained by the same setup, while InternVL-U and CogView4 differ in training data and scale, so the 6.5x parameter gap is a reference point, not a clean architecture control. Main results fix loops at 4; whether the main model keeps gaining beyond 8 loops is only probed on B/32. The looping-beats-denoising-steps claim holds in the 25-to-50-step range and is untested in few-step or distilled regimes. Everything rests on one backbone, so interactions with other architectures go unexamined.