VibeEdit Replaces Text Prompts with Canvas Marks, Scoring 79.9 on Edit Benchmark

Sydney-Uni · hf · 2026-10-09

Sydney Uni's VibeEdit introduces canvas instructions—spatial marks plus short notes placed directly on the image—to specify where and what to edit without separate text prompts. Built on Qwen-Image-Edit with layer-decoupled conditioning, trained on 1.55M edit pairs with region-weighted SFT and rubric-guided RL, it scores 79.9 VLM rubric and 32.8 dB outside-region PSNR on a 419-case human-curated benchmark, versus 67.4 and 24.0 dB for FireRed, the strongest text-instructed baseline.

Original post →

More from Multimodal

Multimodal channel →