VibeEdit: Image Editing with Canvas Instructions
Jinjing Zhao, Fangyun Wei, Yitong Wang, Xiuyu Wu, Yunuo Chen, Yang Yue, Sirui Zhang, Wenbo Wang, Hongyang Zhang, Dong Chen, Yan Lu, Chang Xu
cs.CV, cs.AI
2026-10-09
VibeEdit treats circles, arrows, and handwritten notes drawn on the image as the edit instruction, scoring 79.9 on a 419-case benchmark, 12.5 points above the best text baseline.
Instruction-based editors handle "change the color" fine and fumble on "which of the three identical chairs". The user compresses a spatial reference into prose and the model guesses the referent, so wrong-target edits are the standard failure mode. Pointing tools fix location: DragGAN drags points, MagicQuill paints brushes, but semantics still ride on a separate text prompt, splitting one edit into two inputs. VibeEdit merges them. Circles, scribbles, arrows, and short handwritten notes drawn on the image form a canvas instruction, with no separate text prompt, covering addition, removal, replacement, attribute modification, and movement.
Three parts.
Data: 1.55M pairs with annotations rendered online. Florence-2 proposes candidate objects, GPT-5.6-sol picks one worth editing and writes the edit description, SAM 3 segments the mask. Removal comes from ObjectClear inpainting and reverses into addition; attribute edits and moves come from Qwen-Image-Edit; replacement from FLUX.1 Fill inside the mask. The design choice that matters: storage and rendering are decoupled. The corpus stores images, masks, and structured edit fields, and canvas instructions are rendered on the fly during training with randomized stroke shapes, widths, and handwriting or print typefaces. One edit pair yields many annotation styles, so the model learns the semantics of an instruction instead of one particular hand.
Architecture: layer-decoupled conditioning. Feed the composited annotated image into the generative pathway and strokes leak into the output. VibeEdit splits the input: a frozen Qwen2.5-VL reads the annotated image for semantic tokens, while the VAE separately encodes the clean source image and the instruction layer rasterized on a gray background as DiT conditions. The ablation prices the split: source only 24.7, annotated composite 58.6, two separate layers 67.8.
Training: region-weighted SFT, then rubric-guided RL. In a local edit most pixels do not change, and a uniform loss dilutes the region that does, so object and annotation regions get a 1.5x loss weight over a mask dilated by 50 pixels. RL follows DiffusionNFT: a VLM judge answers binary questions on edit success, outside preservation, and local quality, plus a penalty when outside-region PSNR falls below 25 dB. The 3,520 RL conditions are picked for high reward variance across rollouts, because variance is what provides the comparison signal.
The benchmark has 419 human-curated cases built around picking the right target among similar objects. Text baselines get a human-written instruction averaging 21.3 words; VibeEdit gets the marked image and nothing else. FireRed is the strongest text-instructed baseline.
| Method | Input | VLM rubric | Outside-region PSNR |
| FireRed-Image-Edit-1.0 | detailed text | 67.4 | 24.0 dB |
| GPT-Image-1 | annotated image | 42.7 | 13.4 dB |
| VibeEdit-base (SFT only) | canvas instruction | 67.6 | 32.7 dB |
| VibeEdit (full) | canvas instruction | 79.9 | 32.8 dB |
Numbers worth pulling out:
For editing products, the paper shows the cost of pointing can move from language to canvas, and the model side is a LoRA adaptation of Qwen-Image-Edit rather than a retrain. Circling an object and jotting a word beats writing a spatially precise sentence.
For researchers, two reusable ideas. Storing edits and rendering annotations online applies to any interaction signal (clicks, boxes, trajectories) you want to turn into training data. Layer-decoupled conditioning addresses the general problem of instructions that must guide generation without appearing in the output.
To be clear, the components are off the shelf; the contribution is the interface, the data scale, and the combination. Incremental, but a clean increment.