VibeEdit Replaces Text Prompts with Canvas Marks, Scoring 79.9 on Edit Benchmark
Sydney-Uni · hf · 2026-10-09
Sydney Uni's VibeEdit introduces canvas instructions—spatial marks plus short notes placed directly on the image—to specify where and what to edit without separate text prompts. Built on Qwen-Image-Edit with layer-decoupled conditioning, trained on 1.55M edit pairs with region-weighted SFT and rubric-guided RL, it scores 79.9 VLM rubric and 32.8 dB outside-region PSNR on a 419-case human-curated benchmark, versus 67.4 and 24.0 dB for FireRed, the strongest text-instructed baseline.
More from Multimodal
- TerraVis quantifies world-grounded visual consistency failures in text-to-image models — the-aiml · 2026-10-09
- LMArena's post-training recipe lifts Flux2dev by 69 Elo on T2I leaderboard — lmarena-ai · 2026-10-09
- You can now spot Opus AI video slop by its sound: synced beats as a fingerprint — hudzah · 2026-10-09
- Monkey King riding a tiger: AI video nails a stunning Chinese-style action scene — lucky-plume · 2026-10-09
- Autoregressive Retriever (ARR) Refines Queries with Retrieved Item Feedback via SFT and RL — _reachsumit · 2026-10-09
- Sony's Syn-Omni: Shared + Expert LoRA Paths Beat Omnimodal Embedding Baselines Across 81 Tasks — _reachsumit · 2026-10-09