Appearance Pointers add localized multimodal control to diffusion transformers

Rahul Sajnani · hf · 2026-07-22

Appearance Pointers introduces a modality-agnostic way to control where text or image cues influence a Diffusion Transformer.

The method uses compact pointer tokens aligned with user-specified masks, plus a region correspondence network and spatial aggregation to route appearance cues to the right locations. The authors say it is the first localized multimodal control interface for a DiT that does not require retraining the base model from scratch. Across metrics, the single model matches or exceeds modality-specific state-of-the-art methods for regional control in image synthesis.

Original post →

More from Multimodal

Multimodal channel →