Latent-Identity Tuning for Text-to-Image Face Editing

Yifan Zhong · hf · 2026-07-14

The paper proposes **Latent-Identity Tuning** to solve the issue of "insufficient identity precision" in personalized text-to-image editing, making it particularly suited for fine-grained facial edits. ### Core Idea - Rather than directly editing the input image, it **modifies the latent representation of a specific identity** to generate multiple images with different styles but consistent identity. - The method relies on a **frozen encoder** and requires no additional training; the authors discover controllable identity editing dimensions by exploring semantic directions in the latent space. - These latent tokens correspond to different facial regions or semantic attributes, enabling **local, semantically consistent** modifications. ### Experimental Conclusions - In both qualitative and quantitative experiments, the method achieves various local facial edits while maintaining solid cross-image identity consistency. - The project page is now public.

Related event: Training-Free Latent-Identity Tuning for Text-to-Image Face Editing(2 posts)→

Original post →

More from Multimodal

Multimodal channel →