Appearance Pointers add localized multimodal control to diffusion transformers
Rahul Sajnani · hf · 2026-07-22
Appearance Pointers introduces a modality-agnostic way to control where text or image cues influence a Diffusion Transformer.
The method uses compact pointer tokens aligned with user-specified masks, plus a region correspondence network and spatial aggregation to route appearance cues to the right locations. The authors say it is the first localized multimodal control interface for a DiT that does not require retraining the base model from scratch. Across metrics, the single model matches or exceeds modality-specific state-of-the-art methods for regional control in image synthesis.
More from Multimodal
- Grok Build turns one prompt into a full ARPG with AI-generated game assets — tetsuoai · 2026-07-22
- Grok 4.5 builds a walkable 3D theme park with rides — minchoi · 2026-07-22
- A dad built a controller-ready game in two hours with Grok 4.5 — minchoi · 2026-07-22
- ComfyUI Wan dance test renders a 30-second clip in 70 minutes — tostane · 2026-07-22
- ConsiSpace lifts video spatial reasoning by 12.6 points with geometry-aware memory — Ting Huang · 2026-07-22
- xAI says Grok Imagine will make a full-length Odyssey movie this year — kevinnbass · 2026-07-22