NUS rethinks DiT residual connectivity: 1.73x fewer training iterations, 1.39 FID

NationalUniversityofSingapore · hf · 2026-09-29

Researchers from the National University of Singapore rethink residual connections in Diffusion Transformers (DiTs), which currently integrate all preceding layers into a monolithic residual stream. A systematic analysis of DiT's internal representations reveals a latent preference for early-layer feature reuse and symmetric layer guidance.

They propose a structured connectivity design that turns residuals from passive summation into an active retrieval mechanism: each block selectively "attends" to critical earlier representations through direct, differentiable cross-depth paths, rather than static skip connections or dense all-layer routing.

Results: faster convergence with up to 1.73x fewer training iterations, and significant FID gains with under 0.1% extra parameters — improving REPA-XL/2 from 5.9 to 4.34 FID without guidance, and reaching 1.39 FID with classifier-free guidance.

Original post →

More from Multimodal

Multimodal channel →