DiT study finds template tokens store semantics and enables 20% FLOPs pruning
RTP-LLM · hf · 2026-07-22
Text template tokens act as semantic registers in diffusion transformers
This paper studies how text-to-image diffusion transformers (DiTs) compute during denoising with a causal interpretability framework that decomposes attention and intervenes across token spans, heads, and layers.
- The authors separate prompt-content tokens from structural template tokens and find that template tokens carry little prompt-specific information at the encoder output.
- Surprisingly, those structural tokens become dominant image-to-text attention sinks and causally preserve object identity inside the DiT, functioning as implicit semantic registers.
- The paper argues that semantics are not transferred directly from prompt tokens to template tokens; instead, prompt meaning is first injected into image latents and then read back into template tokens.
- Based on this, the authors propose a training-free pruning rule: heads that attend most strongly to prompt tokens are dispensable.
- Pruning those heads removes 20% of attention FLOPs with only a 1.4-point drop on GenEval.
- The work also maps computation across depth and heads, separating semantic routing from visual synthesis and showing a progression from identity formation to propagation and refinement.
More from Multimodal
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Non-coder builds full-featured Android ComfyUI client with ChatGPT, submits to Google Play — ComfierUI · 2026-09-11
- FastH3-Live hits 22fps: acceleration node benchmarks and the --vram-headroom trick — spartong945 · 2026-09-11
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11