DiT study finds template tokens store semantics and enables 20% FLOPs pruning
RTP-LLM · hf · 2026-07-22
Text template tokens act as semantic registers in diffusion transformers
This paper studies how text-to-image diffusion transformers (DiTs) compute during denoising with a causal interpretability framework that decomposes attention and intervenes across token spans, heads, and layers.
- The authors separate prompt-content tokens from structural template tokens and find that template tokens carry little prompt-specific information at the encoder output.
- Surprisingly, those structural tokens become dominant image-to-text attention sinks and causally preserve object identity inside the DiT, functioning as implicit semantic registers.
- The paper argues that semantics are not transferred directly from prompt tokens to template tokens; instead, prompt meaning is first injected into image latents and then read back into template tokens.
- Based on this, the authors propose a training-free pruning rule: heads that attend most strongly to prompt tokens are dispensable.
- Pruning those heads removes 20% of attention FLOPs with only a 1.4-point drop on GenEval.
- The work also maps computation across depth and heads, separating semantic routing from visual synthesis and showing a progression from identity formation to propagation and refinement.
More from Multimodal
- Dreamina Seedance 2.0 is being tested as a major AI video upgrade — Med1_Ai · 2026-07-22
- Midjourney creator shares four surreal “imaginery tokens” for identity-themed portraits — LudovicCreator · 2026-07-22
- Elon Musk reposts a Grok Imagine prompt for dreamy 8-bit fairy art — elonmusk · 2026-07-22
- Qwen image-edit nodes aim to preserve input resolution and cut VRAM use — Fine-Run992 · 2026-07-22
- Krea 2 Identity Edit returns with new samples and prompts — Fishmongr · 2026-07-22
- xAI’s Grok Imagine showcases sprawling fantasy worlds in a new image drop — elonmusk · 2026-07-22