Qwen Image 2.1 falls back to image editing when given reference images
extra2AB · reddit · 2026-09-24
A Reddit user testing Qwen Image 2.1 found it can't yet work as text-to-image with references like closed models such as Nano Banana: feeding 4-5 reference images makes the model edit the first image instead of generating a new one. They also hit anatomy issues like six-fingered hands, and are asking whether a reference-based T2I workflow exists.
More from Multimodal
- Microsoft claims MAI is hill-climbing 61% faster than rivals, gains +283 Elo on text-to-image in 10 months — SchoeneggerPhil · 2026-09-24
- Real-time vision model runs expression, object detection and finger counting under 1 second — LinusEkenstam · 2026-09-24
- Kling 4.0 leak: 120-second videos, Omni Reference with 15 elements, up to 4K images — koltregaskes · 2026-09-24
- Limestone claims Claude Opus 5.5 one-shotted its entire launch video — alex_verem · 2026-09-24
- Using Ling-3.0-flash-VL to analyze 8 ultrasound reports and solve a pregnancy puzzle — alifcoder · 2026-09-24
- Claude synthesizes a full band from math, noise and its own voice — willcb · 2026-09-24