Sora's Diffusion Transformer explained by hand in 14 steps: prompts enter only as scale and shift

ProfTomYeh · x · 2026-09-16

ProfTomYeh walks through Sora's Diffusion Transformer (DiT) in 14 hand-worked steps. Core takeaway: DiT is a transformer that predicts noise — video enters as latent patches, never pixels, and the prompt and timestep touch the video only via adaptive layer norm, entering as a scale and a shift rather than a separate model.

Key flow:

Steps 4 and 11 are training, 12–14 are generation, and the rest is the block you stack dozens of times. DiT's core idea lives on as the foundation of today's video generation models.

Original post →

More from Multimodal

Multimodal channel →