Qwen-Image-Edit preserves ~90% of reference object identity in inpainting, but takes 4 min per inference
Super_Mission_3130 · reddit · 2026-09-07
A developer details experiments on reference-based object replacement (source image + mask + reference image → same scene with the object swapped), where identity preservation of the reference object is the hard requirement.
- SDXL Inpainting with ControlNet and IP-Adapter only produced visually similar objects, losing structural details.
- Qwen-Image-Edit-2511 with ReCoEdit-RL preserved roughly 90% of the reference object's identity, including details SDXL struggled with.
- The trade-off: high VRAM usage and 4 minutes per inference on their setup, too slow for repeated production use.
The author is asking for open-source pipelines that balance reference fidelity, VRAM and speed — ideally adapting the reference object's geometry/viewpoint first, then transferring appearance with correct scale, perspective, lighting and shadows.
More from Multimodal
- Reddit user shares WIP AI-generated dark fantasy short film 'Wanderers' — DaWid_Shapiro · 2026-09-11
- Creator makes 2D electro-pop anime music video with just a prompt using MiniMax H3 — Hailuo_AI · 2026-09-11
- Street View to driving footage: GPT Astra fetches images, MiniMax H3 turns them into dashcam video — Hailuo_AI · 2026-09-11
- Single-author ECCV 2026 paper makes rolling shutter correction practical — ducha_aiki · 2026-09-11
- AI digital human covers Japanese classic so realistically viewers can't tell — JourneymanChina · 2026-09-11
- ComfyUI Style Explorer Adds LoRA Preview Catalog and Sharing — neonsparksuk · 2026-09-11