EVR Reward Model: Breaking the Consistency Bottleneck in Multi-Reference Image Editing
Yingmao Miao · hf · 2026-08-03
Current image editing models struggle to maintain visual consistency and overall harmony in multi-reference editing. Directly using MLLMs as zero-shot evaluators faces a tension between hallucination-prone long-form reasoning and limited short-form deductive power.
- EVR Method: Proposes the Multi-dimensional Evaluation-Verification Reward (EVR). It decomposes evaluation into distinct visual criteria; an MLLM Evaluator generates multiple candidate hypotheses, and a Verifier grounds each claim in concrete visual evidence to accept or reject it, producing reliable and fine-grained reward signals.
- Performance: Combined with a scalable data pipeline, this method enables RL fine-tuning of off-the-shelf editors without architectural changes. Extensive experiments on the base Qwen-Image-Edit show substantial gains in consistency and harmony, matching or surpassing NanoBanana.
More from Multimodal
- Motion Skill Turns Claude into a Video Team: Generate Launch Videos from a URL — Scobleizer · 2026-08-03
- ComfyUI Adds Day 0 Support for MiniMax Video Model, Slashing VRAM by 66% for RTX 3060 — crystal_alpine · 2026-08-03
- Qwen3.8-Max Takes #2 on Vision Arena, Just Behind Claude — arena · 2026-08-03
- MiniMax-H3 Model Card Surfaces: Omni-modal Generation with Native Stereo Audio — ostrisai · 2026-08-03
- MiniMax-H3 Weights Now Available on Hugging Face — blahblahsnahdah · 2026-08-03
- ComfyUI Newbie Seeks Help: Matching Models with Components — Myrliandre · 2026-08-03