DyRef lifts Qwen-Image-Edit-2511 from 4.97 to 8.38 on multi-reference editing
量子位 · wechat · 2026-07-29
- DyRef targets complex multi-reference image generation, where a model must combine identity, background, pose, lighting, and style from multiple reference images without letting them conflict.
- The team built OmniRef-Bench, a 395-sample benchmark covering 5 reference types, 2 to 7 reference images per sample, and up to 4 simultaneous reference categories.
- DyRef uses a two-stage training recipe: supervised fine-tuning first, then reinforcement learning with dynamic reward optimization so harder samples get more weight.
- On OmniRef-Bench, the post says Qwen-Image-Edit-2511 improved from 4.97 to 8.38 in MLLM evaluation average, narrowing the gap to NanoBananaPro and beating Seedream4.5.
- The article also claims the method improves performance on MultiBanana, OmniContext, and even single-image editing benchmarks such as ImgEdit and DreamBench++.
More from Multimodal
- Describe your dream world to an AI dragon, which generates the planet for you — repligate · 2026-08-24
- Using kintsugi texture to fix cracks in edited 3D meshes — repligate · 2026-08-24
- Generating Hannibal Character Videos with FL2VA Model — Nimblecloud13 · 2026-08-24
- MiniMax H3 Revives Medieval Short Stories: Complete Workflow Shared — zanatas · 2026-08-24
- NAPE Audio Pretraining Achieves SOTA Without Decoders — kastnerkyle · 2026-08-24
- H3 excels at generating complex space scenes — SIR_NVAX_A_LOT · 2026-08-24