DiT study finds template tokens store semantics and enables 20% FLOPs pruning

RTP-LLM · hf · 2026-07-22

Text template tokens act as semantic registers in diffusion transformers

This paper studies how text-to-image diffusion transformers (DiTs) compute during denoising with a causal interpretability framework that decomposes attention and intervenes across token spans, heads, and layers.

Original post →

More from Multimodal

Multimodal channel →