CityU Hong Kong unveils RefineEdit, a training-free image editing framework that tops PIE-Bench

CityU-HongKong · hf · 2026-09-21

Researchers at City University of Hong Kong introduce RefineEdit, a training-free prompt-to-prompt image editing framework built on a Generative Refinement Network.

Key idea: couple edit localization with content generation by globally refining binary image codes. An editing branch is initialized from an intermediate source state (reusing the emerging layout), and the signed probability differences between the two branches over the same source-sampled bits select which positions and bits to edit — selected bits follow editing refinement while the rest copy the evolving source state.

Stabilization mechanisms:

Results: across nine editing categories on PIE-Bench, RefineEdit achieves the best background-preservation scores (PSNR, LPIPS, MSE, SSIM) plus the highest whole-image and edited-region CLIP scores, addressing both incomplete edits and accidental alteration of unrelated regions that plague diffusion and causal autoregressive editors.

Original post →

More from Multimodal

Multimodal channel →