NUS rethinks DiT residual connectivity: 1.73x fewer training iterations, 1.39 FID
NationalUniversityofSingapore · hf · 2026-09-29
Researchers from the National University of Singapore rethink residual connections in Diffusion Transformers (DiTs), which currently integrate all preceding layers into a monolithic residual stream. A systematic analysis of DiT's internal representations reveals a latent preference for early-layer feature reuse and symmetric layer guidance.
They propose a structured connectivity design that turns residuals from passive summation into an active retrieval mechanism: each block selectively "attends" to critical earlier representations through direct, differentiable cross-depth paths, rather than static skip connections or dense all-layer routing.
Results: faster convergence with up to 1.73x fewer training iterations, and significant FID gains with under 0.1% extra parameters — improving REPA-XL/2 from 5.9 to 4.34 FID without guidance, and reaching 1.39 FID with classifier-free guidance.
More from Multimodal
- German children's media firm Blue Ocean hires Visual AI Artist fluent in ComfyUI, German required — Additional-Log-2617 · 2026-09-29
- Same prompt showdown: Opus 5.5 Max effort vs Grok 4.7 Fast xhigh for motion graphics video — FinanceYF5 · 2026-09-29
- ComfyUI beginner asks how to separate pose, identity and background references — butterfly--1313 · 2026-09-29
- LoRA fixes dark, dirty outputs from Qwen image model, author shares on Civitai — Comfortable-Mind4875 · 2026-09-29
- Indie dev builds a Manus-style AI promo video for CellMotion — yihui_indie · 2026-09-29
- Kling 4.0 launches with major lipsync and 15-image Omni Reference, first music video hands-on — tomchapin · 2026-09-29